Antigravity vs Cursor vs Codex vs Claude: Which AI Coding App I Actually Keep Open

I’ve spent the last stretch of side-project work switching between four AI coding apps: Google Antigravity, Cursor, OpenAI Codex, and Claude. They all promise roughly the same thing: describe what you want, and an agent reads your code, edits files, runs commands, and hands you a working change.

The demos all look great. Living with them is different.

This is not a benchmark. I didn’t run a test suite across models or time tasks with a stopwatch. These are my personal impressions from building real projects, and they’re colored by the kind of work I do: lots of small, independent apps, lots of TypeScript, and a growing amount of AI agent code. Your mileage will vary, and these products change every few weeks.

TL;DR

AppMy takeBest for
AntigravityThe least ergonomic of the four. Too much setup and too many approvals, and it leans on models I don’t love for agentic work.Trying out Google’s take on an agent-first IDE
CursorSolid editor experience, but my Claude usage ran out fast and the fallback models weren’t as good.People who want an editor first and an agent second
CodexExcellent for building AI agents. I haven’t used it heavily for general coding yet because I kept hitting limits.Agent and LLM-app development
ClaudeThe one I enjoy most and get the most value from, mostly thanks to Opus 5.5.Day-to-day agentic coding across a whole repo

What I care about

Before getting into each app, here’s what decides whether a tool stays in my workflow:

  1. Ergonomics. How much friction sits between “I have an idea” and “the agent is working on it”? Setup, configuration, and approval prompts all count.
  2. Model quality for agentic work. Autocomplete is easy now. What matters is whether the model can plan a multi-file change, run the tests, read the failure, and fix it without me babysitting.
  3. Usage limits. A great model you can only use for an hour isn’t that great. I want to finish a session without being pushed to a worse model halfway through a task.
  4. Speed. Not raw tokens per second. I mean how long it takes to get to a correct result.

Antigravity: the most friction

Antigravity was the worst of the four for me, ergonomically.

The first problem was setup. It needed more configuration than the others before I felt productive, and once it was running, I spent a surprising amount of time approving scripts and permissions. Some approval flow is good. I want an agent to ask before it does something destructive. But there’s a balance, and Antigravity’s prompts were frequent enough that I couldn’t step away and let it work. An agent that needs me to click “allow” every minute isn’t saving me much attention.

The second problem was the models. Antigravity puts the Gemini Flash models front and center. Flash is fast and cheap, and that’s useful for plenty of things. For agentic coding, where the model has to hold a plan across many steps and recover from its own mistakes, it didn’t hold up against the models I use elsewhere.

The third problem followed from the first two: it felt slower to actually get things done. Between approvals, extra setup, and more back-and-forth to correct the model, tasks took longer end to end, even when individual responses came back quickly.

I want to be fair here: Antigravity is newer than the others, and Google moves fast. I’d revisit it. As of today, it’s the one I reach for least.

Cursor: good, until the good model runs out

Cursor is a clear step up. It feels like a polished editor first, and the AI features are well integrated into it. If you live in VS Code, it’s the easiest of the four to adopt.

My problem with Cursor was usage. I prefer using Claude models inside Cursor, and I burned through my Claude usage fast. Once that happened, I was effectively pushed onto other models. In practice that meant Grok, and for the kind of work I was doing, it wasn’t a good substitute. The change in quality mid-session was jarring: a task that was going well would start going sideways after the switch.

That’s the core tension with a multi-model editor. Being able to pick any model is a feature. But if the model you actually want is the one with the tightest limit, the flexibility mostly means you get downgraded.

Codex: amazing for building agents

Codex surprised me in a specific area: building AI agents. When I was working on agent code, things like tool definitions, orchestration loops, and prompt plumbing, Codex was genuinely excellent. It seemed to understand the shape of that kind of software well, and it produced code I’d keep.

I haven’t used Codex much for general coding yet, and the reason is simple: I kept running into limits. Every time I started leaning on it, I’d hit a wall and have to switch tools. That’s a shame, because what I did use was strong. I’ll keep it in rotation for agent-heavy projects and give it a fairer shot at everyday coding once the limits stop getting in the way.

Claude: the one I keep open

Claude is the one I ultimately enjoy the most, and it isn’t especially close.

Most of that comes down to Opus 5.5. It’s the model where I most often hand off a real task, something like “add this feature, update the tests, make the build pass,” and come back to something correct. It plans well, it reads the codebase before guessing, and when a test fails it actually works out why instead of flailing. That’s the difference between a tool I supervise and a tool I delegate to.

The ergonomics are good too. It works directly in my repo, runs my real commands, and the permission model asks about the things that matter without nagging me about everything else. That’s the opposite of my Antigravity experience. I can give it a well-scoped task, go do something else, and come back to progress instead of a pending approval dialog.

I get a lot of value out of it every day, and it’s the tool most of my side projects are being built with right now.

How I’d choose

If you’re picking one today, here’s my honest advice:

The best thing you can do is test them on your work. Take a real task from your backlog, something that touches a few files and has tests, and give it to each tool. Within an afternoon you’ll know which one you trust.

For me, that test ended with Claude.