Which AI coding agent supports what? Every cell cited.
AgentMatrix compares 22 AI coding agent CLIs across 32 capabilities. Each of the 704 cells stores the sentence from the vendor's own documentation that justifies it, the URL of that page, and the date the sentence was last found there. A free GitHub Action re-checks every quote weekly.
The problem
"Does Codex CLI have hooks? Can Gemini CLI run as an MCP server? Which of these agents run natively on Windows?" In 2026 these questions come up daily, and they are usually answered from memory. The comparison tables that exist are rarely dated, almost never sourced, and go stale within weeks because Claude Code, Codex CLI, Gemini CLI, Cursor, GitHub Copilot CLI, OpenCode, Cline, Goose, Aider and Amp ship features constantly.
AgentMatrix answers from the documentation and shows its work. It is deliberately narrower than some hand-maintained comparisons such as coding-agents-matrix, which cover more products. The difference is what backs each cell.
What is different
- Every value has a receipt. A cell is only shown as yes, partial or no when its quote was found, verbatim, at the linked URL. Formatting differences are tolerated (whitespace, capitalisation, punctuation, markdown syntax), wording differences are not.
- Quotes are re-checked every week, for free. A GitHub Action re-fetches all source pages and searches for every quote. A quote that has disappeared demotes its cell to unknown and opens an issue. The table can go stale into "unknown"; it cannot silently drift into a wrong "yes" when a documentation page is rewritten.
- The decision rule is written down. Each capability has a question and a rubric that says exactly what counts as yes, partial or no (for example: a Docker image offered purely as an install method is not a sandbox; a feature labelled experimental is partial even when it is on by default). A disagreement about a cell is a disagreement about the rubric, not about taste.
- Cells are derived by Claude, then verified mechanically. When documentation changes, a maintainer or contributor runs the pipeline with their own API key: Claude reads the agent's documentation in one request, decides every capability against the rubric, and returns a quote per cell. Only quotes that survive the mechanical check are kept, and a human reviews the resulting pull request row by row.
What the matrix covers
Agents (10): Claude Code (Anthropic), Codex CLI (OpenAI), Gemini CLI (Google), Cursor CLI, GitHub Copilot CLI, Cline, OpenCode, Goose (Block), Aider, Amp (Sourcegraph).
Capabilities (32), in six groups:
What the data says (September 2026)
A few rows that stood out while building it. Every claim below links to the same evidence the matrix shows; open the live matrix and click the cell to read the quote.
| Capability | Finding |
|---|---|
| Acts as an MCP server | Yes only for Claude Code (claude mcp serve). Codex CLI still documents a deprecated stdio MCP server and Goose can re-expose its built-in extensions; the other seven do not offer it, and the ones that do expose themselves to other tools use ACP or HTTP, which is not MCP. |
| Local models | Yes for six (Codex CLI, GitHub Copilot CLI, Cline, OpenCode, Goose, Aider). No for Claude Code, Cursor CLI, Amp, and for Gemini CLI, whose "local model" is an experimental Gemma classifier that only routes requests; every answer still comes from hosted Gemini. |
| Open source | Six of ten are OSI-licensed (Codex CLI, Gemini CLI, Cline, OpenCode, Goose, Aider). Claude Code, Cursor CLI, GitHub Copilot CLI and Amp are proprietary. |
| Free tier | Yes for Codex CLI, Gemini CLI, Cursor CLI and GitHub Copilot CLI. Partial for Cline, OpenCode, Goose and Aider (rotating free models or free only through a third-party provider). No for Claude Code (requires a paid plan) and Amp (its free tier is closed to new sign-ups). |
| Built-in sandbox | Yes for Claude Code, Codex CLI, Gemini CLI and Cursor CLI. Partial for GitHub Copilot CLI (preview), Goose (containers you set up) and Amp (remote machines). No for Cline, OpenCode and Aider; a Docker image as an install method does not count. |
| Checkpoints and rewind | Six agents rewind their own edits without git. Codex CLI advises the user to make git checkpoints instead, Goose has no rewind, Aider can undo its last commit, Amp exposes an undo tool. |
| Native Windows | Nine of ten run natively; Amp supports Windows through WSL only. |
| AGENTS.md | Seven read it automatically. Claude Code reads CLAUDE.md and documents an import workaround, Gemini CLI needs the filename configured, Aider has no instruction-file convention. |
| Persistent memory written by the agent | Yes only for Claude Code's auto memory. Five are partial (preview, experimental, off by default, or a documented methodology), four have only a user-maintained file. |
| Universal | All ten have a full-auto mode, a headless or CI mode, model selection, reasoning-effort control and image input. |
How it works
Two loops, one free and one that costs API credits.
Every week, free
1. fetch every evidence URL in data/matrix.json (raw markdown or page text)
2. search each page for the cell's quote, tolerating formatting but not wording
3. quote found -> verified_at = today
quote missing -> value = "unknown"; the old quote, URL and value stay in the notes
page unreachable -> cell untouched, reported
4. no value changed -> the refreshed dates are committed
values changed -> a pull request and an issue list the affected cells for a human
When documentation changes, with an API key
1. fetch every source page listed for the agent in data/agents.json
2. Claude reads all of it in one request and returns one JSON entry per capability:
value, verbatim quote, source URL, notes, constrained by a schema
3. every returned quote is searched for in the fetched text; a quote that is not found
demotes the cell to "unknown" (a fabricated or paraphrased quote does not survive)
4. if the new answer is "unknown" but the previous verified quote is still on its page,
the previous cell is kept, so one bad answer never erases good data
5. the result is diffed against the previous matrix and opened as a pull request
The mechanical check proves that a sentence exists on a page. Whether the sentence supports the value is a judgment made by the rubric, the model and the reviewers, in that order. Cells are open to challenge: report one with a link to the docs.
What it costs
Nothing to run. GitHub Actions and GitHub Pages are free for public repositories, and the weekly quote check needs no API key. Re-deriving cells with Claude is paid by whoever runs it, with their own key: at list prices roughly $0.50 to $3 per agent with claude-fable-5-1, on the order of $30 for a full pass over all 22 agents. The repository ships no key, and it does not need one to stay honest.
Contributing
More agents are welcome. Adding one means listing its documentation URLs in data/agents.json, filling the 32 cells (with your own key, by hand, or leaving them unknown for someone else), and running npm run check, which fails if any quote is not found at its URL. Capabilities are open for proposals too; a good proposal comes with a rubric two people would apply the same way.
- Repository: github.com/barbarkaragul-oss/agentmatrix (MIT)
- Live matrix: barbarkaragul-oss.github.io/agentmatrix
- Raw data: matrix.json, with a
CITATION.cffin the repo for citing it
Türkçe özet
AgentMatrix, 22 yapay zekâ kodlama ajanı komut satırı aracını (Claude Code, Codex CLI, Gemini CLI, Cursor CLI, GitHub Copilot CLI, Cline, OpenCode, Goose, Aider, Amp, Qwen Code, Crush ve 10 tane daha) 32 yetenek üzerinden karşılaştıran açık kaynak bir özellik matrisi. Her hücre, üreticinin kendi dokümanından birebir bir alıntı, o sayfanın adresi ve alıntının son doğrulandığı tarihle birlikte saklanıyor; ücretsiz bir GitHub Action her hafta bütün alıntıları kaynağında yeniden kontrol ediyor, kaybolan alıntının hücresi "bilinmiyor"a düşüyor ve bir issue açılıyor. Çalıştırması ücretsiz, lisansı MIT, verisi JSON olarak indirilebilir.