Eight weeks of user evidence across Claude Code, Codex, Cursor, GitHub Copilot and Gemini reveal a new competitive axis: not who writes the best code—but who can be trusted with the longest leash.
The last eight weeks were not a simple model horse race. Every major coding system gained power. The surprise was what users started complaining about once that power arrived.
Cloud agents, nested subagents, automatic routing and million-token models expanded what developers could delegate. But the conversation moved just as quickly toward context loss, runaway usage, invisible quotas, session failures and output that looked right while being wrong.
The durable advantage is shifting from model capability to operational trust: durable state, observable execution, predictable cost, stable releases and safe rollback.
“Users cannot reliably predict the outcome, cost or collateral behavior of a long-running coding-agent task.”
THE TOKEN REPORT, BASE-CASE THESIS
5systems tracked
39TruthAPI calls
6,000ranked slots returned
2,024unique records
THE TOKEN REPORTCLAUDE CODE
01 · CURRENT LEADER, HIGHEST VOLATILITY
Claude Code: power without a governor
9-MONTH CALLRELATIVE TRACTION AT RISKMEDIUM CONFIDENCE
The harness is the advantage. TruthAPI connects Claude Code to subagents, skills, MCP, plan-execute workflows, memory and long-horizon editing—not merely a strong model.
Practitioners still notice the difference.
“Claude Code is best because it has dynamic workflows.”
CLAUDE CODE DISCORD USER · JUL 10
WHY TRACTION IS AT RISK
Failure can look like success. A fresh issue reports a clean code review after subagents had failed at the spend limit—while billing continued.
Scale magnifies context loss. One current task spawned four subagents; eight compactions per agent produced 40 reported off-track incidents. A new rate-limit issue adds service risk.
UNSOLVED PAINPredictable control of autonomous execution
Bound the agents, tokens, permissions and behavioral drift—without neutering the system.
THE CALL: Claude remains top-tier. Stability-first releases, hard spend controls and reliable rollback could erase this risk.
The switcher signal repeats. Moves from Claude to Codex appeared in both fixed evidence windows, alongside current claims that Codex is the better implementation engineer.
The model line moved again. GPT-5.6 shipped across Codex and the API, while practitioners compared OpenAI's usage resets favorably with Anthropic's.
2 windowsshowing migration signals—not one selectively chosen launch week
WHAT COULD STOP IT
The harness still trails. Current practitioner evidence says the underlying model is competitive while workflows and subagents remain less mature than Claude Code.
Product state is not durable yet. Fresh issues report Windows freezing when changing projects and a broken agent-thread picker.
UNSOLVED PAINDurable long-horizon context
Keep tools, task state and history intact across sessions, compaction, platforms and updates.
THE CALL: Codex gains relative practitioner traction if OpenAI fixes product-state debt before it compounds.
THE TOKEN REPORTCURSOR
03 · THE HARNESS BECOMES A PLATFORM
Cursor: the integrated-team bet
9-MONTH CALLGAIN LIKELYMEDIUM CONFIDENCE
JUN 29 mobile praiseJUL 8 workflow-trained modelJUL 11 ACP reaches editorsJUL 19 routing complaint
WHY THE HARNESS COMPOUNDS
Cursor is becoming portable. ACP exposes its agent inside Zed and IntelliJ instead of confining it to one editor.
Workflow data may become a moat. Grok 4.5 was described as trained with Cursor data from real developer workflows. Mobile and fast-edit preferences broaden the surface.
Cursor's edge is increasingly the workflow—not a captive model.
THE TOKEN REPORT · SYNTHESIS
WHAT COULD STOP IT
Routing opacity erodes trust. A current report says Cursor switched from Grok to Sonnet without intent and consumed API usage.
Parallel agents lack shared state. Fresh evidence describes Claude Code and Cursor silently contradicting each other on the same repository.
UNSOLVED PAINPredictable usage economics
Translate models, cache, fast mode and tokens into productive hours and a believable bill.
THE CALL: Cursor gains if its workflow layer stays model-agnostic and makes routing and billing inspectable.
THE TOKEN REPORTGITHUB COPILOT
04 · DISTRIBUTION MEETS AGENT MODE
Copilot: the default, not yet the favorite
9-MONTH CALLENTERPRISE GAIN · POWER USERS FLATMEDIUM CONFIDENCE
The agent now reaches the work queue. Copilot can accept a Jira ticket, stream progress, change code and open a pull request without leaving Jira.
GitHub owns the execution surface. Its cloud agent can research, plan, branch, test and iterate inside a GitHub Actions environment.
20%reported code-review cost reduction after rewriting instructions—with no measured quality loss
WHY PREFERENCE LAGS
Advanced users still separate it from long-horizon agents. Current comparisons describe Copilot mostly as autocomplete and weaker on multi-step context.
A local win did not generalize. GitHub says the focused instructions that improved code review did not improve Copilot CLI's broader exploratory work.
Show users what routing chose, what a task will cost and why agent mode is better than completion.
THE CALL: Enterprise use rises through workflow adjacency. Broad practitioner preference waits on a coherent agent experience.
THE TOKEN REPORTGEMINI / ANTIGRAVITY
05 · THE BEAUTIFUL WILDCARD
Gemini: specialist upside, trust deficit
9-MONTH CALLSPECIALIST GAIN · GENERALIST UNCERTAINLOW CONFIDENCE
JUN 18 Flash speed signalJUL 3–6 coding failuresJUL 15 working game shipsJUL 16 Flash preference
WHY SPECIALIST UPSIDE IS REAL
The technical surface is substantial. Gemini 3.5 Flash documents a 1M-token input window, code execution, search grounding and fast multi-step loops. Antigravity adds a managed sandbox.
A non-programmer shipped a working artifact.
Hundreds of prompts produced a playable browser game in one 2 MB HTML file.
R/GEMINIAI USER · JUL 15
WHY GENERALIST TRUST IS WEAK
Current users report fragile coding work. One migrated from Antigravity to Claude after repeated errors; another hit a limit after roughly 30 minutes and switched providers.
Approval boundaries regressed. A current report says Antigravity changed files without asking where Gemini CLI and Claude Code required approval.
UNSOLVED PAINGrounded correctness
Preserve the speed and visual imagination; remove fabricated repository facts and unsafe execution.
THE CALL: Gemini gains in rapid loops, UI and prototyping. Generalist leadership waits on approval, state and quota reliability.
THE TOKEN REPORTTHE FORECAST · LEADERS
BASE CASE · THROUGH APRIL 2027
The call—and the evidence that could break it
Relative practitioner traction among serious active users. Direction is qualitative; confidence is editorial—not a measured probability.
CLAUDE CODE
↘ RELATIVE RISKMEDIUM CONFIDENCE
WHY
The deepest agent-workflow harness in the graph is also accumulating spend, compaction and reliability failures. The pattern appears in both fixed evidence windows.
First-party harnesses can narrow the moat. Opaque model switching and usage charging could turn orchestration into a liability.
THE TOKEN REPORTTHE FORECAST · CHALLENGERS
BASE CASE · THROUGH APRIL 2027
Distribution versus specialist velocity
GITHUB COPILOT
→ ENTERPRISE GAIN POWER USERS FLATMEDIUM CONFIDENCE
WHY
Copilot now reaches the issue queue, repository, Actions and pull request. Distribution should lift enterprise use; current advanced-user comparisons still favor other long-horizon agents.
If GitHub turns issue, code, CI, policy and review into one coherent loop, distribution becomes product advantage and power-user preference could rise sharply.
GEMINI / ANTIGRAVITY
↗ SPECIALIST GAIN GENERALIST UNCERTAINLOW CONFIDENCE
WHY
Fast loops, 1M-token context, code execution and a managed sandbox support prototyping. Current quota, approval and coding-reliability reports weaken the generalist case.
Google has speed, context, grounding and distribution. A unified developer surface with reliable approval and quotas could create broad gains quickly.
WHAT COUNTS
Direction, then calibration
Issue 002 will score these calls with the same symmetric queries and consecutive fixed windows. Numeric probabilities return only after the calls earn calibration.
THE DECIDING VARIABLE
Operational trust
Capability is abundant. Durable state, bounded execution, predictable cost, observable work and safe rollback remain scarce.
THE TOKEN REPORTPRACTITIONER PLAYBOOK
DON’T PICK A RELIGION. DESIGN A PORTFOLIO.
The stack for the next quarter
PRIMARY IMPLEMENTER
Codex or Claude Code
Use Codex where session stability is proven; keep Claude for deep reasoning and orchestration with explicit agent and spend caps.
INTEGRATED TEAM WORK
Cursor
Best fit for codebase context, local-to-cloud handoff, background agents and pull-request flow. Track spend outside the product.
ENTERPRISE DEFAULT
GitHub Copilot
Exploit repository and policy integration. Benchmark routed models and agent mode against a simpler completion baseline.
VISUAL SPECIALIST
Gemini
Delegate UI exploration and graphical apps. Never accept product facts, APIs or architecture without independent verification.
FIVE RULES FOR OPERATORS
Set a token and wall-clock budget before launch.Stop conditions are part of the prompt, not an afterthought.
Separate planning, implementation and review models.Cross-model review catches shared harness blind spots.
Persist state in the repository.Assume the next session remembers nothing.
Require evidence-bearing completion.Tests, diffs, logs and uncertainty—not “done.”
Measure the system, not the demo.Time-to-verified-output, repairs, regressions and total cost.
THE TOKEN REPORTUSE-CASE RADAR
WHAT BUILDERS ARE DOING NOW
Where coding agents are actually going
The common use case is no mystery. The important signal is what happens when builders connect code execution to a measurable outcome.
ESTABLISHED USEREPEATEDacross source types
Bounded implementation
Give the agent a concrete change and an objective check it can run.
FEATUREBuild against acceptance tests.
BUGFix a reproducible failure.
REFACTORChange structure while behavior stays green.
Method: verification loop Problem defeated: false completion
EMERGINGRISINGmultiple demonstrations
Closed-loop computational research
The agent proposes, runs, measures and revises—not merely writes the research code.
BRAIN2QWERTY
An Auto Research workflow found word-error-rate improvements beyond conventional hyperparameter optimization.
SOCIAL SCIENCE
Frontier coding agents reproduced computational findings as reliable workflow executors.
Signal: strongest for Claude Code, with additional Codex research evidence.
FRINGE, BUT REALEARLYone to three cases
Coding escapes software
GENEALOGYStructured family-history and archival auto-research.
POKERSolver bots used to test reasoning and optimization.
EDGE CAMERASAn MCP server running directly on an AXIS IP camera.
Pattern: these differ by system. The interface determines the edge.
THE BIGGER SHIFTCode is becoming the agent’s universal actuator.
The product is no longer always software. Sometimes software is simply how the agent reaches the outcome.
The current evidence window runs from May 25 through July 20, 2026; the comparison window is March 30 through May 24. All five products received the same progress, outcomes, friction and competition scans, then a prior-window scan, a graph traversal and full-record inspection. Waves were used only as a cross-check.
THE FRESH EVIDENCE RUNCoverage is not prevalence
39TruthAPI calls, all reply times preserved
6,000current-window product result slots
2,024exact-ID-deduplicated current records
617event records in the product cohort
758connected object records
649technical documentation passages
THE EVIDENCE UNIT
One exact-ID-deduplicated TruthAPI record: an event, an object or a documentation passage. It is not a unique person, company, vote or controlled observation.
WHAT CARRIED WEIGHT
First-party documentation, direct artifacts, GitHub issues, reproducible outcomes, repeated practitioner reports and signals that survived both fixed windows.
WHAT DID NOT
Raw mention volume, a single viral post, bundled seats, product marketing or rank alone. Counts describe retrieval coverage—not adoption or market share.
SELECTION DISCIPLINE
Each product received 1,200 current result slots and 300 prior-window slots. Exact IDs were deduplicated before candidate selection. `getSet` inspected the decisive records in full.
LIMITATIONS
This is directional qualitative evidence, not a representative survey. Source communities overlap. User reports can conflate model and harness quality. Discord and Reddit signals are self-reported. Direction bands express editorial judgment.
WHAT COMES NEXT
Issue 002 will score every call using the same protocol. Numeric probabilities return only after repeated forecasts earn calibration.
REPRODUCIBILITY
The 39 queries, full return payloads, selected records, analysis and every reply time are preserved in research/token-report-refresh-2026-07-20/.
The research archive preserves all 39 query arguments, full return payloads, exact reply times, deduplication metrics, selected records and editorial analysis. Counts describe returned evidence—not unique people, audited market share or controlled trials.