Julio
Performance marketing + engineering
Buys and runs every tool himself against the fixed task set before it gets a number.
Method v1.2 · published 4 Jan 2026 · last amended 12 Jun 2026
Every tool in this category is put through the same published rubric and the same fixed set of twelve tasks, run against three real repositories. Nothing here is a feature grid copied from a vendor’s site — every number comes from a hands-on run we paid for ourselves.
The six scoring criteria and the exact percentage each one carries are published in the rubric and are fixed for the whole category. We do not re-tune them per tool, and we do not author a headline anywhere — the score you see is derived from the per-criterion results, so the panel and its bars can never drift apart.
The testing is done by Julio and the aitools.reviews team. We do not invent reviewer bylines, and we do not publish a test log we did not run. If a byline is on a verdict, that person ran the tasks.
We make money three ways — affiliate rev-share, newsletter sponsorship, and a paid review queue where a vendor can buy a testing date but never a verdict. All three are disclosed on the page. Money buys attention and timing; it does not buy a score, and we do not remove a verdict to keep a partner happy.
Scores carry a date and a version. We re-test on a thirty-day cycle, and when a score moves we log the change with the reason. The method itself is versioned too — see the changelog below for what changed and when.
| Criterion | Weight | Measured by |
|---|---|---|
| Task completion | 25% | Tasks finished with no human edit |
| Code quality | 20% | Blind diff review by a second engineer |
| Speed | 15% | Wall-clock time per completed task |
| Value for money | 15% | Total cost of the 12-task run |
| Learning curve | 15% | Time to first useful output, cold start |
| Team & security | 10% | SSO, retention policy, training opt-out |
We do not remove or rewrite a published verdict to please a vendor. Every correction is dated and logged in the open, and the old score stays visible next to the new one.
Anonymous reviews are how directories lose their credibility. Ours are named.
Julio
Performance marketing + engineering
Buys and runs every tool himself against the fixed task set before it gets a number.
aitools.reviews team
Hands-on testing
The people behind aitools.reviews who run each tool through the same tasks and publish the full log.