Method v1.2 · published 4 Jan 2026 · last amended 12 Jun 2026

How we test AI coding agents

Every tool in this category is put through the same published rubric and the same fixed set of twelve tasks, run against three real repositories. Nothing here is a feature grid copied from a vendor’s site — every number comes from a hands-on run we paid for ourselves.

The six scoring criteria and the exact percentage each one carries are published in the rubric and are fixed for the whole category. We do not re-tune them per tool, and we do not author a headline anywhere — the score you see is derived from the per-criterion results, so the panel and its bars can never drift apart.

Who does the testing

The testing is done by Julio and the aitools.reviews team. We do not invent reviewer bylines, and we do not publish a test log we did not run. If a byline is on a verdict, that person ran the tasks.

Where the money comes from

We make money three ways — affiliate rev-share, newsletter sponsorship, and a paid review queue where a vendor can buy a testing date but never a verdict. All three are disclosed on the page. Money buys attention and timing; it does not buy a score, and we do not remove a verdict to keep a partner happy.

When we re-test

Scores carry a date and a version. We re-test on a thirty-day cycle, and when a score moves we log the change with the reason. The method itself is versioned too — see the changelog below for what changed and when.

The rubric — coding agents

Weighted rubric for AI coding agents
CriterionWeightMeasured by
Task completion25%Tasks finished with no human edit
Code quality20%Blind diff review by a second engineer
Speed15%Wall-clock time per completed task
Value for money15%Total cost of the 12-task run
Learning curve15%Time to first useful output, cold start
Team & security10%SSO, retention policy, training opt-out

The protocol

  1. 1. Buy the tool ourselves — every plan, on our own card, never a vendor comp.
  2. 2. Run the same twelve tasks across three real repositories, with no cherry-picking.
  3. 3. Score each criterion against the published rubric, blind-reviewed by a second engineer.
  4. 4. Record the price we paid that day and snapshot every plan into the ledger.
  5. 5. Re-test on a fixed thirty-day cycle and log any score that moves.
  6. 6. Publish with the date, version and full test log, so anyone can check our work.

Where the money comes from

  • Affiliate rev-share on outbound links — disclosed above every CTA, and it never touches the score.
  • Newsletter sponsorship on "the weekly verdict" — clearly labelled, kept separate from editorial.
  • A paid review queue — vendors can buy a testing date, never a verdict and never a score.

We do not remove or rewrite a published verdict to please a vendor. Every correction is dated and logged in the open, and the old score stays visible next to the new one.

Who does the testing

Anonymous reviews are how directories lose their credibility. Ours are named.

Julio

Performance marketing + engineering

Buys and runs every tool himself against the fixed task set before it gets a number.

aitools.reviews team

Hands-on testing

The people behind aitools.reviews who run each tool through the same tasks and publish the full log.

v1.2 · 12 Jun 2026
Raised Team & security from 5% to 10% and re-scored every affected tool against the new weighting.
v1.1 · 2 Mar 2026
Added the cross-service migration task to the fixed set of twelve.
v1.0 · 4 Jan 2026
First published method and rubric for AI coding agents.