ChatGPT Dots vs Claude vs Gemini: Which AI Agent Actually Fits Your Work?

17분 읽기조회 0
공유

Imagine hiring an assistant who writes beautifully, answers instantly, and confidently announces that the project is finished. Then you open the spreadsheet. The totals are wrong, the source links are missing, and the customer email has already been sent.

You did not hire an assistant. You acquired a second job.

That is the useful starting point for comparing the latest AI agents. Not which one produces the most impressive demonstration, but which one leaves you with less work—and fewer surprises.

This comparison focuses on three ecosystems: ChatGPT Dots and Work/Codex, Claude’s Cowork experience and Claude Code, and Gemini Enterprise with Google’s developer tooling. They overlap, but they are not interchangeable products. This is a source-based buying guide as of October 4, 2026, not a hands-on benchmark or an exhaustive ranking of every agent on the market. The scenarios and recommendations below are editorial analysis.

Three companies, three different ideas of an assistant

OpenAI’s Dots make ongoing delegation the central idea. A dot runs in the cloud, has its own computer and browser, and can keep working while your computer is off. It uses GPT-6 Astra. Think of the product proposition as “keep this responsibility moving,” rather than merely “answer this message.” Official Dots guide

Anthropic is reducing the distinction between asking and delegating. Its September announcement merges chat and Cowork into one Claude experience, initially rolling out to Pro and Max. You describe the outcome; the interface increasingly handles whether that requires a reply or a task. Claude’s product announcement

Google’s Gemini Enterprise is more naturally understood as a workplace platform: connected business information, shared agents, and centrally managed access. It connects to systems including Google Workspace and Microsoft 365, with edition-dependent support for custom and third-party agents. Gemini Enterprise overview

My shorthand is a persistent assistant, a conversational workbench, and an organizational agent platform. Those are starting points for evaluation, not exclusive capabilities. Claude can handle recurring work; OpenAI has enterprise collaboration; Google is not limited to searching documents.

The mistake is buying one because its homepage sounds like another.

Give them a launch, not a trivia question

Suppose a small software company is shipping a new feature. It needs a competitor brief, a landing-page update, a customer announcement, and a weekly adoption report.

That apparently simple assignment contains three different jobs.

The first is making something. Can the agent turn rough notes into a useful brief or document? Does it preserve important qualifications? Can you revise one section without damaging the rest? Claude’s combined conversation-and-task experience is an obvious candidate to evaluate here, but polished prose alone does not establish a winner.

The second is maintaining something. Can it notice a changed launch date, keep a report current, and bring back only the decisions that need attention? Dots is explicitly positioned around this continuity. The test is whether it reduces reminders you must write, not whether it sends more notifications.

The third is coordinating access. Which customer records may the agent read? Can a teammate reuse the workflow without inheriting somebody else’s permissions? What happens when an employee leaves? Gemini Enterprise belongs on this shortlist when shared data and centrally managed agents are the problem. That does not make it automatically the best purchase for a solo writer.

Before picking a vendor, identify which of these jobs is actually consuming your day. Otherwise, you may pay for organizational governance when you need a good drafting partner—or buy a personal assistant when your real problem is fragmented company data.

Coding is a separate contest

A good office agent is not automatically the best coding agent. For repository work, compare Codex, Claude Code, and Google Antigravity in the actual development environment you intend to use. Google describes Antigravity as an agentic development ecosystem, with a developer SDK in preview and a Google Cloud integration path. Availability depends on the offering; an ecosystem announcement is not a promise that every seat has every tool. Antigravity’s official overview

The new model names matter, but they are not product test results. GPT-6.1 Sol and Astra, Claude Sonnet 5.5 and Opus 5.5, and Gemini 4 Argon describe model choices or announced capabilities—not a standardized race between complete applications.

There is a concrete reason to be cautious. Google’s Argon evaluation document says that some competitor results come from provider reports. Its OSWorld comparison omits Anthropic results because the reported subsets differ. It also describes different evaluation setups across tasks. A single combined “agent intelligence score” would conceal those differences. Google’s evaluation methodology

For coding, I would judge a finished patch on four things: did the tests pass, did behavior remain correct, did the agent stay within scope, and how much human review was needed?

A patch that passes tests after changing fifteen unrelated files may be less useful than a smaller, slower patch. Likewise, an agent that says it could not run the test suite is more useful than one that gives you the impression that everything passed.

Do not pay for a stronger model to compensate for a broken development environment. Missing dependencies, inaccessible test data, and unclear acceptance criteria can make any candidate look bad.

The price comparison has two receipts

The first receipt is the subscription. The second is the work you still have to do.

Here are useful entry points—not equivalent bundles. USD list prices were checked on October 4; taxes, billing terms, region, and rollout can change what you actually pay or receive.

Product routePublished price referenceWhat the number does not tell you
ChatGPT Work/Codex through Plus$20/monthIt is not the entry price for Dots.
ChatGPT Pro, with eligible Dots access$100, $200, or $500/monthHigher price does not mean unlimited execution.
Claude Pro / MaxPro: $20 monthly; Max starts at $100/monthIncluded usage and model access still matter.
Gemini Enterprise BusinessStarts at $21/seat/monthA business seat is not equivalent to a personal agent plan.
Gemini Enterprise Standard / PlusCombined listing starts at $30/seat/monthThis is not a confirmed flat price for every Plus configuration.

Sources: ChatGPT pricing, Dots eligibility, Claude pricing, Google’s edition pricing.

Dots access is rolling out with plan, age, region, and workspace conditions. Its conversations do not consume ChatGPT usage limits, but tasks it starts or manages in Work or Codex consume those products’ allowances. That distinction matters more than the phrase “always on.” Dots access and usage

API prices are a different receipt again. Do not divide a subscription fee by a model’s token price and claim that the result is the number of tasks included. OpenAI’s pricing documentation explicitly separates API prices from subscription usage. Subscription usage guidance

Now add your time. Suppose Agent A costs $30 a month but creates three extra hours of checking, while Agent B costs $100 and creates thirty minutes. At a hypothetical $40 per hour, the totals are $150 and $120. Those are invented numbers, not measured product results. Their purpose is to expose the expense that pricing tables omit.

The better metric is cost per accepted outcome, including review and rework. An inexpensive agent is not inexpensive if you have to reconstruct its reasoning from twenty browser tabs.

“Works in the background” deserves a literal test

Close the laptop. Change a requirement. Let a login expire. See what happens.

Those tests reveal more than watching another perfect demo. An agent may be able to run in the cloud while still depending on a particular local file, browser session, or connected application. “Cloud-based” does not make every dependency cloud-based.

A timely example: Anthropic’s documentation says new Cowork tasks on Pro and Max will run in the cloud from October 6, 2026, while existing local tasks remain local. That is a scheduled change, not something this October 4 article treats as already completed. Local-folder access still has its own desktop dependency. Cowork cloud transition

Continuity also has a less glamorous side: stopping. Can you pause the job without losing its useful results? Can you understand what remains unfinished? Can somebody else pick it up?

For recurring work, a graceful interruption is a feature. Silence, duplicated actions, or a falsely cheerful “done” are operational failures even if the model is excellent at reasoning.

More connections can mean more cleanup

Connector counts are a poor purchasing shortcut. A connection might support reading records but not updating them, or work differently under an organization’s permissions. Inspect the specific action you need, not just the familiar logo.

For each candidate, I would test whether it can cite the exact source, distinguish a draft from a committed change, and recover when an integration fails. I would also ask how to export useful work and remove access when leaving the product.

Privacy deserves the same specificity. “Enterprise security” is not a substitute for checking retention, training use, local-device access, auditability, and the relevant edition. A personal account and an enterprise deployment should not inherit each other’s promises by association.

There is a practical rule here: connect enough to complete one valuable workflow, then expand. Connecting everything on day one makes both success and failure harder to explain.

What I would shortlist—and what would change my mind

For an individual who repeatedly forgets follow-ups, I would evaluate Dots first. Its continuous-delegation design matches the problem. I would reject it if notification noise and review effort outweighed the reminders it removed.

For a writer, analyst, or small team moving between discussion and deliverables, I would start with Claude and ChatGPT Work on the same files. I would compare editability, source quality, and how well each handles corrections—not the confidence of its opening response.

For a company whose information is scattered across managed systems, I would include Gemini Enterprise alongside the business offerings from OpenAI and Anthropic. Identity, data access, and workflow ownership would drive the decision before writing style.

For a developer, I would run a paid pilot of Codex, Claude Code, or Antigravity on a few completed issues with known acceptance criteria. Start from the same repository state, keep permissions comparable, record retries, and count human minutes. If you intentionally test different models or budgets, label the result as a comparison of configurations, not pure model intelligence.

These are evaluation priorities, not measured winners. The answer changes with your tools, access, workload, and tolerance for intervention.

The latest agents are becoming better at doing things. Choosing one still requires a less exciting question: after it finishes, what is left for you?

If the answer is “approve a clear result,” you may have found an assistant. If it is “investigate what happened,” keep shopping.


한국어 번역

글을 잘 쓰고, 답도 빠르며, 프로젝트가 끝났다고 자신 있게 말하는 비서를 채용했다고 상상해보자. 그런데 스프레드시트를 열어보니 합계가 틀렸고, 출처 링크가 없고, 고객 이메일은 이미 발송됐다.

비서를 채용한 것이 아니다. 일이 하나 더 생긴 것이다.

최신 AI 에이전트를 비교할 때 유용한 출발점은 여기에 있다. 누가 가장 인상적인 데모를 보여주는지가 아니라, 누가 내 일과 예상치 못한 문제를 줄여주는가다.

이 글은 ChatGPT Dots와 Work/Codex, Claude의 Cowork 경험과 Claude Code, Gemini Enterprise와 구글 개발 도구라는 세 생태계를 비교한다. 기능이 겹치지만 서로 같은 제품은 아니다. 2026년 10월 4일 공식 자료를 바탕으로 작성한 선택 가이드이며, 직접 수행한 벤치마크나 시장의 모든 에이전트를 망라한 순위표는 아니다. 아래 시나리오와 추천은 필자의 분석이다.

세 회사가 생각하는 비서는 서로 다르다

OpenAI Dots의 중심은 지속적인 업무 위임이다. 자체 컴퓨터와 브라우저를 갖고 클라우드에서 실행되며, 사용자 컴퓨터가 꺼져 있어도 일을 이어갈 수 있다. 기반 모델은 GPT-6 Astra다. “이 메시지에 답해줘”보다 “이 일을 계속 챙겨줘”에 가까운 제품이다. Dots 공식 안내

Anthropic은 질문과 업무 위임의 구분을 줄이고 있다. 9월 발표는 채팅과 Cowork를 하나의 Claude 경험으로 합치며, Pro와 Max부터 순차 적용한다. 사용자가 원하는 결과를 말하면, 단순 답변인지 실행이 필요한 업무인지 인터페이스가 처리하는 방향이다. Claude 제품 발표

Gemini Enterprise는 업무 플랫폼으로 이해하는 편이 자연스럽다. 회사 정보 연결, 공유 에이전트, 중앙 접근 관리가 중심이다. Google Workspace와 Microsoft 365 등을 연결하며, 자체·외부 에이전트 지원은 에디션에 따라 다르다. Gemini Enterprise 소개

나는 이를 계속 챙기는 비서, 대화형 작업 공간, 조직용 에이전트 플랫폼으로 요약한다. 평가의 출발점이지 독점 기능을 뜻하지는 않는다. Claude도 반복 업무를 처리하고, OpenAI도 기업 협업을 지원하며, 구글도 문서 검색만 하는 것은 아니다.

문제는 한 제품의 홈페이지가 다른 제품처럼 들린다는 이유로 구매하는 것이다.

상식 퀴즈 대신 출시 업무를 맡겨보자

작은 소프트웨어 회사가 새 기능을 출시한다고 가정하자. 경쟁사 조사, 랜딩 페이지 수정, 고객 공지, 주간 사용 현황 보고서가 필요하다.

간단해 보이는 이 요청에는 서로 다른 세 가지 일이 들어 있다.

첫째는 무언가를 만드는 일이다. 거친 메모를 쓸 만한 보고서나 문서로 바꾸는가? 중요한 단서를 보존하는가? 한 부분을 고칠 때 나머지를 망가뜨리지 않는가? 대화와 실행을 합친 Claude는 여기서 평가할 자연스러운 후보지만, 매끄러운 문장만으로 승자가 정해지지는 않는다.

둘째는 무언가를 유지하는 일이다. 출시일 변경을 반영하고 보고서를 최신으로 유지하며, 사람의 판단이 필요한 것만 가져오는가? Dots는 이런 지속성을 명시적으로 겨냥한다. 알림을 더 많이 보내는지가 아니라, 내가 다시 써야 하는 지시를 줄이는지가 시험 기준이다.

셋째는 접근을 조율하는 일이다. 어떤 고객 기록을 읽어도 되는가? 동료가 다른 사람의 권한까지 물려받지 않고 업무 흐름을 재사용할 수 있는가? 직원이 퇴사하면 어떻게 되는가? 공유 데이터와 중앙 관리가 문제라면 Gemini Enterprise를 후보에 넣을 이유가 있다. 그렇다고 혼자 글을 쓰는 사람에게 자동으로 최선의 구매가 되지는 않는다.

먼저 이 셋 중 무엇이 하루를 잡아먹는지 구분하자. 그렇지 않으면 좋은 초안 작성 도구가 필요한데 조직 관리 기능에 돈을 쓰거나, 흩어진 회사 데이터가 문제인데 개인 비서를 사게 된다.

코딩은 별도의 경기다

좋은 사무 에이전트가 곧 최고의 코딩 에이전트는 아니다. 저장소 작업이라면 실제 사용할 개발 환경에서 Codex, Claude Code, Google Antigravity를 비교해야 한다. 구글은 Antigravity를 에이전트형 개발 생태계로 소개하며, 프리뷰 SDK와 Google Cloud 연결 경로를 안내한다. 제공 범위는 상품별로 다르다. 생태계 발표가 모든 좌석에서 모든 도구를 쓸 수 있다는 약속은 아니다. Antigravity 공식 소개

새 모델 이름은 중요하지만 제품 전체의 시험 결과는 아니다. GPT-6.1 Sol과 Astra, Claude Sonnet 5.5와 Opus 5.5, Gemini 4 Argon은 모델 선택지나 발표된 역량을 뜻한다. 완성된 앱들을 동일 조건에서 겨룬 결과가 아니다.

주의할 구체적인 이유도 있다. 구글의 Argon 평가 문서는 경쟁사 점수 일부를 해당 회사의 발표에서 가져왔다고 설명한다. OSWorld에서는 평가 부분집합이 달라 Anthropic 점수를 제외했다. 과제마다 실행 조건도 다르다. 이를 하나의 ‘에이전트 지능 점수’로 합치면 차이가 가려진다. 구글 평가 방법론

코딩 결과물은 네 가지로 평가하고 싶다. 테스트가 통과했는가, 기존 동작이 유지됐는가, 요청 범위를 지켰는가, 사람이 얼마나 검토해야 했는가.

무관한 파일 열다섯 개를 바꾼 뒤 테스트를 통과한 패치는, 작고 느린 패치보다 덜 유용할 수 있다. 테스트를 못 돌렸다고 밝히는 에이전트도 모든 검증이 끝난 것처럼 보이게 하는 에이전트보다 유용하다.

망가진 개발 환경을 더 비싼 모델로 보완하려고 하지 말자. 의존성 누락, 접근 불가능한 테스트 데이터, 모호한 완료 기준은 어떤 후보든 나쁘게 보이게 만든다.

가격 비교에는 영수증이 두 장 있다

첫 번째는 구독료다. 두 번째는 여전히 내가 해야 하는 일이다.

아래는 진입 가격을 이해하기 위한 표이며 동일 구성의 상품 비교가 아니다. 10월 4일 확인한 미국 달러 표시 가격으로, 세금·결제 조건·지역·순차 제공 여부에 따라 실제 가격과 기능이 달라질 수 있다.

이용 경로공개 가격 기준숫자만으로 알 수 없는 것
Plus의 ChatGPT Work/Codex월 $20Dots의 진입 가격은 아니다.
Dots 이용 자격이 있는 ChatGPT Pro월 $100·$200·$500비싸다고 실행이 무제한은 아니다.
Claude Pro / MaxPro 월 결제 $20, Max 월 $100부터제공량과 모델 접근 조건이 중요하다.
Gemini Enterprise Business좌석당 월 $21부터기업 좌석은 개인 에이전트 요금제와 다르다.
Gemini Enterprise Standard / Plus통합 안내상 좌석당 월 $30부터모든 Plus 구성의 확정 단일 가격은 아니다.

출처: ChatGPT 가격, Dots 이용 자격, Claude 가격, 구글 에디션 가격.

Dots는 요금제·나이·지역·워크스페이스 조건에 따라 순차 제공된다. dot과의 대화는 ChatGPT 한도를 소모하지 않지만, dot이 시작하거나 관리하는 Work·Codex 작업은 해당 제품의 제공량을 사용한다. ‘상시 실행’이라는 표현보다 중요한 구분이다. Dots 접근과 사용량

API 가격은 또 다른 영수증이다. 구독료를 모델의 토큰 단가로 나눈 뒤 그 결과를 포함 작업 수라고 제시해서는 안 된다. OpenAI 가격 문서도 API 요금과 구독 사용량을 구분한다. 구독 사용량 안내

이제 내 시간을 더해보자. 에이전트 A는 월 $30이지만 추가 검토에 세 시간이 들고, B는 $100이지만 30분이 든다고 가정하자. 시간의 가치를 시간당 $40으로 잡으면 총비용은 각각 $150과 $120이다. 실제 제품 측정값이 아닌 가상의 숫자다. 가격표에 빠진 지출을 드러내기 위한 예시다.

더 좋은 지표는 검토와 재작업을 포함한 채택 가능한 결과물 하나당 비용이다. 브라우저 탭 스무 개를 뒤져 작업 과정을 재구성해야 한다면 싼 에이전트도 싸지 않다.

‘백그라운드에서 작동’은 말 그대로 시험해야 한다

노트북을 닫아보자. 요구사항을 바꿔보자. 로그인이 만료되면 어떻게 되는지 보자.

이런 시험은 완벽한 데모를 한 번 더 보는 것보다 많은 것을 알려준다. 에이전트가 클라우드에서 실행돼도 특정 로컬 파일, 브라우저 세션, 연결 앱에 의존할 수 있다. 실행 장소가 클라우드라고 모든 의존성이 클라우드로 옮겨지는 것은 아니다.

시의성 있는 예도 있다. Anthropic 문서는 2026년 10월 6일부터 Pro·Max의 새 Cowork 작업이 클라우드에서 실행되고 기존 로컬 작업은 그대로 남는다고 안내한다. 10월 4일 기준인 이 글에서 이미 완료된 변화로 취급하지 않는다. 로컬 폴더 접근에도 별도의 데스크톱 의존성이 있다. Cowork 클라우드 전환 안내

지속성에는 덜 화려한 측면도 있다. 멈추는 능력이다. 유용한 결과를 잃지 않고 중단할 수 있는가? 무엇이 미완료인지 알 수 있는가? 다른 사람이 이어받을 수 있는가?

반복 업무에서는 정상적인 중단도 기능이다. 침묵, 중복 실행, 실제와 다른 ‘완료했습니다’는 모델의 추론 능력이 뛰어나도 운영 실패다.

연결이 많아지면 정리할 일도 늘어날 수 있다

커넥터 개수는 구매를 결정하기에 좋지 않은 지름길이다. 기록을 읽을 수 있어도 수정은 못 할 수 있고, 조직 권한에 따라 다르게 작동할 수 있다. 익숙한 로고보다 필요한 구체적 동작을 확인해야 한다.

각 후보가 정확한 출처를 제시하는지, 초안과 실제 반영을 구분하는지, 연동 실패에서 복구하는지 확인하고 싶다. 제품을 떠날 때 쓸 만한 결과물을 내보내고 접근을 해제하는 방법도 살펴볼 것이다.

개인정보 보호도 구체적으로 봐야 한다. ‘기업용 보안’이라는 표현이 보관 정책, 학습 활용, 로컬 접근, 감사 가능 여부, 해당 에디션 확인을 대신하지 못한다. 개인 계정과 기업 배포가 서로의 약속을 자동으로 공유하는 것은 아니다.

실용적인 원칙은 하나다. 가치 있는 업무 하나에 필요한 만큼 연결한 다음 확대하자. 첫날부터 모든 것을 연결하면 성공과 실패의 이유를 모두 설명하기 어려워진다.

내가 먼저 검토할 후보와, 생각을 바꿀 조건

후속 일을 자꾸 놓치는 개인이라면 Dots부터 평가하겠다. 지속적인 위임이라는 설계가 문제와 맞기 때문이다. 다만 사라진 재촉보다 알림과 검토 부담이 더 커진다면 선택하지 않겠다.

논의와 결과물 제작을 오가는 작가·분석가·소규모 팀이라면 같은 파일로 Claude와 ChatGPT Work를 비교하겠다. 첫 답변의 자신감보다 수정 편의성, 출처 품질, 피드백 반영을 보겠다.

관리되는 여러 시스템에 정보가 흩어진 회사라면 OpenAI·Anthropic의 기업용 상품과 함께 Gemini Enterprise를 후보에 넣겠다. 문체보다 계정과 권한, 데이터 접근, 업무 소유권이 우선이다.

개발자라면 완료 기준이 알려진 과거 이슈 몇 개로 Codex, Claude Code, Antigravity 중 후보를 유료 시험하겠다. 같은 저장소 상태에서 시작하고 권한을 비슷하게 맞추며, 재시도와 사람의 시간을 기록하겠다. 모델이나 예산을 의도적으로 다르게 했다면 순수 지능이 아니라 구성별 비교라고 표시해야 한다.

이는 평가 우선순위이지 측정으로 가린 승자가 아니다. 사용 도구, 접근 권한, 업무, 개입 허용 수준에 따라 답은 달라진다.

최신 에이전트들은 일을 수행하는 능력이 좋아지고 있다. 그래도 선택할 때는 덜 화려한 질문이 필요하다. 그것이 일을 마친 뒤 내게 무엇이 남는가?

‘명확한 결과를 승인하는 일’이라면 비서를 찾았을 수 있다. ‘무슨 일이 일어났는지 수사하는 일’이라면 더 찾아보자. exec-a7cf69ae-cd6b-4eb2-b7f6-6f5af4f296cc.png

댓글을 작성하려면로그인이 필요합니다.