How we test
Methodology v0.2. Published before results, so you can judge the rules, not just the outcomes.
Measured results
Objective metrics recorded by our test harness (same inputs for every tool). Raw outputs are archived as evidence.
- Identical prompts, inputs and settings for every tool
- Repeated runs to capture consistency
- Wall-clock time, cost per task, failure/refusal rate
- Automatic checks: tests passed, word error rate, text-rendering accuracy, etc.
- Raw outputs archived with hashes as evidence
Editorial judgment
Subjective assessment by our evaluators, blind where possible. This is opinion, labelled as such — not a measured fact.
- Quality rated 1–10 on a published rubric
- Multiple independent evaluators, blind where possible
- Disagreement is reported, not hidden
- Always labelled as opinion — never presented as fact
Core rules
- Rank by job, not in the abstract. Each job has its own task set and its own weights.
- Scores are 0–100 per job, computed from dimension scores × published weights. There is no single global "AI score".
- NOT YET TESTED is shown wherever no completed test run exists. We never estimate, extrapolate or borrow scores.
- Freshness. Every result shows its test date and the product/model version and plan used. Fast-moving categories are retested more often; stale results are flagged.
- Commercial independence. Affiliate status is not a scoring input. See independence & disclosure.
Test suites v1 (launch categories)
Published before testing starts. Measured, panel and editorial criteria are kept separate.
| Category | Inputs | Measured (objective) | Panel (blind, ≥3 raters, 1–10) | Editorial |
|---|---|---|---|---|
| Video | 8 prompts: cinematic shot; 9:16 product ad; image-to-video product shot; human hands/motion; physics; on-screen text; dialogue with native audio; 2-shot character consistency. First output scored (no cherry-picking); seeds/settings logged. Avatar sub-suite: 60-second script, 2 languages. | time-to-output; cost per clip at plan price; max resolution/duration; watermark on tier; refusal/failure rate; commercial-use terms (ToS link) | prompt adherence; visual quality; motion realism; consistency; audio/lip-sync | editing controls; UX; best-for summary |
| Voice | 10 scripts: narration; dialogue; numbers/dates/acronyms; proper names; emotional read; long-form 5 min; 3 languages. Cloning tested only with a consenting team member's voice. | round-trip WER (synthesize → transcribe with a fixed STT model); pronunciation errors; latency (API); price per 1M characters; license terms | naturalness (MOS); clone similarity (ABX); expressiveness | workflow/editor UX |
| Meeting notes | 3 scripted recordings with ground truth: 4-speaker video call (30 min); in-person 2-speaker with background noise; accented/bilingual call. Seeded facts + 8 action items each. | WER; speaker attribution accuracy; action-item recall/precision; hallucinated facts in summary; time-to-notes; bot vs bot-free; data-retention/training policy; price per seat | summary usefulness | integrations; privacy posture |
| Image | 12 prompts: photoreal portrait; product on white; poster with exact text; multi-object spatial prompt; brand style; edit/inpaint a supplied photo; 3-image character consistency; logo/vector. | exact-text accuracy (% characters correct); checklist adherence (% required elements present); time; cost per image; max resolution; rights | aesthetics; realism; edit fidelity | controls/UX |
| Presentations | 3 briefs: 10-slide seed pitch deck from an outline; sales deck from a 2-page document; training deck from notes. | time-to-first-deck; factual errors introduced; slide-count fidelity; PPTX/Google Slides export defects | design quality; narrative clarity | editability; brand-kit support |
Integrity rules: same inputs and settings across tools; date, plan tier and visible model/version recorded; raw outputs published; one re-test per tool per major release; panel raters are blind to tool identity; affiliate status is hidden from raters and attached only after scoring.
Full Testing Lab methodology
Owner: Benchmark Design & Data · Version 0.2 · 2026-10-08 · Machine-readable companions: lab/jobs.json (weights), lab/benchmarks/tasks.json (task registry), data/schema.sql (storage).
v0.2 change log (2026-10-08): added a 12th lab dimension, privacy_data (Privacy & data handling; §3, §3a), weighted 15% for meeting-notes/transcription jobs and voice cloning and 0–4% elsewhere (§4). Every job score is still NOT_YET_TESTED, so no published score moves. v0.1 weights remain in git history.
Status today: no tool has been tested. Every job score is NOT_YET_TESTED. This document defines how scores will be produced. Nothing on the site may show a score that did not come out of this process.
1. Principles
- No fabricated results. A score exists only if an archived output, a measurement or a rating sits behind it. Untested means
NOT YET TESTED, never an estimate. - Measured ≠ judged. Measured results are produced by code or instruments: latency, WER, tests passed, JSON validity, OCR accuracy, cost. Editorial / panel judgments are human rubric ratings. They are stored in separate tables (
test_outputs.metrics_jsonvsevaluations) and shown separately on every page. - Rankings are per job, not per tool. "Best video AI" does not exist; "best for a 15-second product ad" does. Weights are category-specific and adjusted per job.
- Money never touches scores. Affiliate data lives in
affiliate_programs. No scoring code, view or export reads it (v_public_rankingshas no affiliate columns), and raters never see it. See §10. - Reproducible and dated. Every number carries its date, plan tier, model/version string, prompt hash and output hash.
2. Scope of round 1 (locked lineup)
39 tool slots across the five launch categories, as listed in research/mvp-decision.md §3 and mirrored in data/seed/first_round.py and tools.first_round=1:
- Video (12): Google Veo/Flow, Kling, Seedance, Runway, Hailuo, Luma, Adobe Firefly, Grok Imagine, Higgsfield (tested as a platform), plus HeyGen, Synthesia and Colossyan (avatar).
- Voice (6): ElevenLabs, Fish Audio, Cartesia, Gemini TTS (Google), Murf, Speechify. Hume (Octave) was dropped on 2026-10-08 because its TTS API shuts down on 2026-11-13; it is still tracked with status
pivoted. - Meeting notes & transcription (7): Otter, Fireflies, Fathom, Granola, Read AI, tl;dv, and Zoom's built-in assistant (branded ZoomMate on zoom.com as of 2026-10-08).
- Image (8): GPT Image, Gemini Nano Banana Pro, Midjourney, Ideogram, FLUX, Adobe Firefly, Seedream, Recraft.
- Presentations (6): Gamma (now Gamma 5), Beautiful.ai, Canva, Genspark, Copilot in PowerPoint, Plus AI.
All other tracked tools (78 in total in data/tools.db) appear as NOT YET TESTED with facts only. Discontinued tools (Sora) have status='discontinued'; the site hides them from category listings and never tests them. Chatbots, coding and music benchmarks are fully specified but are not launch categories.
3. Dimensions
Twelve lab dimensions, each scored 1–10 using the anchors in lab/benchmarks/_common.md and the category files:
| Dimension | Meaning | Typical source |
|---|---|---|
| quality | how good the output is to a skilled human | judged (blind panel) |
| accuracy | correctness and adherence: facts, prompt adherence, WER, tests passed, exact text | mostly measured |
| speed | time to usable output (category-specific anchors) | measured |
| ease_of_use | steps and time to complete a standard workflow; non-expert success | measured (timed) + judged |
| features | presence of the capabilities the job needs (stems, SCORM, lip-sync, bot-free capture…) | measured (checklist from official docs + verified in test) |
| price_value | effective cost per usable output (§7) | measured |
| reliability | failure, timeout and refusal rates across repetitions and time windows | measured |
| customization | style, brand, voice and format control | measured + judged |
| commercial_rights | licence terms on the tested plan (quoted, with URL) | documented fact → fixed anchor table |
| support | documented support channels/SLA; response to one standard ticket | measured |
| integrations | exports, API and connectors the job needs; export fidelity | measured |
| privacy_data | data handling: training on user data, retention and deletion, raw-media deletion, SOC 2 / GDPR, participant notice or clone consent, bot-free capture (§3a) | documented fact from official policy/trust pages → fixed rubric, verified in test where possible |
User satisfaction is a separate signal (verified user reviews, §6). It is never mixed into the lab dimensions.
3a. privacy_data rubric (v0.2)
The score comes from a desk check of official privacy-policy, trust/security, DPA and pricing pages, run as STT-12 for meeting notes and VOX-06 for voice cloning. Every item is stored in privacy_facts with its URL, verbatim quote, snapshot path and date. Wherever possible, items are then verified in the test account: deletion is actually executed, the opt-out toggle is present, and bot-free capture works.
| # | Item | Points |
|---|---|---|
| a | Training on customer content: not used by default 2.5 · opt-out available 1.5 · used with no opt-out / not stated 0 | 2.5 |
| b | Retention & deletion: configurable retention / auto-deletion 2 · deletion on request only 1 · not stated 0 | 2 |
| c | Raw audio/media deleted after processing (or a zero-data-retention option) | 1 |
| d | Independent audit: SOC 2 Type II 1.5 · Type I 0.75 · ISO 27001 counts as Type I | 1.5 |
| e | GDPR: DPA available and/or EU data residency | 1 |
| f | Transparency to others: participant notification (meeting notes) / consent verification for cloned voices (voice) | 1 |
| g | Capture control: bot-free capture option (meeting notes) / watermarking or detection of generated audio (voice) | 1 |
- Score = max(1, 10 × points ÷ 10). An item counts only if an official page states it. "Not stated" scores 0 and is shown as not stated, never guessed. Enterprise-only items are counted for the plan we test only if that plan is the tested plan; otherwise they are listed as "Enterprise only".
- The dimension is documented-fact based, not marketing-based. A claim that the test contradicts (e.g. deletion leaves the transcript retrievable) sets that item to 0 and is published as a finding.
- Until the desk check for a tool is complete, its facts are shown on the tool page as facts, and no privacy_data score exists. The job score stays
NOT_YET_TESTEDand weight coverage reflects the missing dimension. - Privacy facts never use affiliate data, vendor questionnaires we cannot cite, or third-party summaries.
4. Weights
Category defaults come first; jobs then shift weights with explicit deltas (lab/build_jobs.py) and are renormalised to 100%. The tables below are generated by python3 lab/build_jobs.py --markdown. Do not edit them by hand.
Category default weights (lab score, sums to 100%; methodology v0.2)
| Category / Job | Qual | Acc | Spd | Ease | Feat | Value | Rel | Cust | Rights | Sup | Integ | Priv |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| chatbots | 19 | 24 | 8 | 7 | 8 | 10 | 8 | 4 | 2 | 2 | 6 | 3 |
| video | 30 | 15 | 7 | 7 | 10 | 10 | 6 | 5 | 6 | 2 | 2 | 1 |
| image | 28 | 18 | 6 | 7 | 8 | 10 | 5 | 7 | 7 | 2 | 2 | 1 |
| music | 30 | 15 | 4 | 7 | 8 | 10 | 4 | 7 | 12 | 1 | 2 | 0 |
| coding | 15 | 29 | 8 | 6 | 8 | 10 | 8 | 5 | 1 | 2 | 7 | 2 |
| voice | 29 | 14 | 10 | 6 | 7 | 10 | 6 | 6 | 6 | 2 | 2 | 4 |
| presentations | 24 | 15 | 8 | 14 | 8 | 10 | 5 | 7 | 0 | 2 | 6 | 2 |
| transcription | 8 | 30 | 8 | 6 | 7 | 10 | 5 | 2 | 0 | 2 | 6 | 15 |
Job weights: chatbots
| Category / Job | Qual | Acc | Spd | Ease | Feat | Value | Rel | Cust | Rights | Sup | Integ | Priv |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
llm.research_cited Research with verifiable citations |
16 | 34 | 5 | 7 | 8 | 10 | 8 | 2 | 0 | 2 | 6 | 3 |
llm.business_writing Business writing & editing |
29 | 16 | 10 | 10 | 6 | 10 | 6 | 4 | 2 | 1 | 4 | 3 |
llm.reasoning_analysis Reasoning, math & data analysis |
14 | 39 | 8 | 4 | 8 | 10 | 8 | 2 | 0 | 2 | 3 | 3 |
llm.long_document Long-document Q&A (contracts, reports) |
14 | 34 | 5 | 7 | 10 | 10 | 8 | 2 | 0 | 2 | 6 | 3 |
llm.multilingual_he Hebrew & multilingual work |
27 | 26 | 5 | 7 | 7 | 10 | 8 | 2 | 2 | 2 | 2 | 3 |
llm.everyday_free Best free everyday assistant |
19 | 18 | 8 | 12 | 8 | 24 | 6 | 1 | 0 | 0 | 2 | 3 |
Job weights: video
| Category / Job | Qual | Acc | Spd | Ease | Feat | Value | Rel | Cust | Rights | Sup | Integ | Priv |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
video.cinematic_character Cinematic scene with a consistent character |
38 | 17 | 4 | 3 | 10 | 6 | 6 | 9 | 6 | 1 | 0 | 1 |
video.social_ad_15s 15-second product ad for Instagram/TikTok |
22 | 15 | 12 | 11 | 10 | 12 | 4 | 2 | 10 | 1 | 1 | 1 |
video.avatar_explainer Talking-avatar explainer / training video |
20 | 20 | 7 | 12 | 14 | 8 | 6 | 2 | 4 | 2 | 5 | 1 |
video.image_to_video Animate a product photo (image-to-video) |
32 | 21 | 7 | 7 | 7 | 10 | 6 | 3 | 6 | 1 | 0 | 1 |
video.budget_creator Best value video AI for solo creators |
22 | 15 | 7 | 10 | 8 | 25 | 4 | 2 | 6 | 1 | 0 | 1 |
Job weights: image
| Category / Job | Qual | Acc | Spd | Ease | Feat | Value | Rel | Cust | Rights | Sup | Integ | Priv |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
image.photoreal_product Photoreal product & lifestyle shots |
32 | 20 | 4 | 4 | 8 | 10 | 5 | 4 | 10 | 2 | 0 | 1 |
image.typography_poster Posters, ads & images with correct text |
23 | 30 | 4 | 7 | 8 | 10 | 3 | 4 | 7 | 2 | 2 | 1 |
image.consistent_character Consistent character / brand style across images |
24 | 22 | 3 | 4 | 8 | 10 | 5 | 15 | 7 | 2 | 0 | 1 |
image.photo_editing AI photo editing (inpaint, remove, extend) |
20 | 24 | 6 | 7 | 14 | 8 | 5 | 5 | 4 | 2 | 5 | 1 |
image.vector_logo_icons Vector graphics, icons & logo exploration |
23 | 18 | 3 | 7 | 16 | 9 | 3 | 10 | 7 | 2 | 2 | 1 |
Job weights: music
| Category / Job | Qual | Acc | Spd | Ease | Feat | Value | Rel | Cust | Rights | Sup | Integ | Priv |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
music.full_song_vocals Full song with vocals from lyrics |
38 | 17 | 2 | 7 | 8 | 10 | 4 | 5 | 8 | 1 | 0 | 0 |
music.background_bgm Royalty-free background music for videos/podcasts |
20 | 13 | 4 | 9 | 8 | 13 | 4 | 4 | 22 | 1 | 2 | 0 |
music.instrumental_score Instrumental score / soundtrack to brief |
32 | 21 | 2 | 3 | 8 | 6 | 4 | 11 | 10 | 1 | 2 | 0 |
music.editing_stems Editing, extending & stems for producers |
22 | 12 | 4 | 3 | 20 | 6 | 4 | 12 | 12 | 1 | 4 | 0 |
Job weights: coding
| Category / Job | Qual | Acc | Spd | Ease | Feat | Value | Rel | Cust | Rights | Sup | Integ | Priv |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
code.agent_existing_repo Autonomous issue-to-PR in an existing repo |
10 | 37 | 5 | 3 | 10 | 7 | 12 | 5 | 1 | 2 | 7 | 2 |
code.inline_completion Inline completion & in-editor help |
12 | 24 | 18 | 6 | 5 | 10 | 6 | 5 | 1 | 2 | 10 | 2 |
code.app_builder_nocode Build a working web app without coding |
20 | 22 | 5 | 18 | 8 | 10 | 8 | 2 | 1 | 2 | 4 | 2 |
code.debugging Debugging & fixing failing tests |
10 | 39 | 8 | 6 | 5 | 10 | 8 | 5 | 1 | 2 | 5 | 2 |
code.code_review Code review & security findings |
12 | 37 | 4 | 4 | 8 | 8 | 8 | 5 | 1 | 2 | 10 | 2 |
Job weights: voice
| Category / Job | Qual | Acc | Spd | Ease | Feat | Value | Rel | Cust | Rights | Sup | Integ | Priv |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
voice.narration_longform Long-form narration (audiobook, e-learning) |
36 | 14 | 3 | 6 | 7 | 10 | 9 | 4 | 6 | 2 | 0 | 4 |
voice.realtime_agent Real-time voice agent (low latency) |
17 | 14 | 24 | 3 | 7 | 10 | 10 | 2 | 3 | 2 | 5 | 4 |
voice.voice_clone Cloning a voice (with consent) |
26 | 21 | 3 | 5 | 4 | 6 | 5 | 8 | 5 | 2 | 0 | 15 |
voice.multilingual_dubbing Multilingual voiceover & dubbing |
26 | 22 | 4 | 6 | 11 | 10 | 6 | 3 | 6 | 2 | 2 | 4 |
voice.short_ads Short ad / social voiceover |
29 | 11 | 4 | 10 | 7 | 12 | 4 | 6 | 12 | 2 | 2 | 4 |
Job weights: presentations
| Category / Job | Qual | Acc | Spd | Ease | Feat | Value | Rel | Cust | Rights | Sup | Integ | Priv |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
pres.pitch_deck Investor pitch deck from a brief |
32 | 17 | 5 | 14 | 8 | 9 | 3 | 7 | 0 | 2 | 2 | 2 |
pres.doc_to_deck Turn a long report into a deck |
20 | 26 | 8 | 10 | 8 | 10 | 5 | 4 | 0 | 2 | 6 | 2 |
pres.on_brand_office On-brand slides inside PowerPoint / Google Slides |
19 | 15 | 5 | 10 | 8 | 7 | 5 | 13 | 0 | 2 | 16 | 2 |
pres.teacher_lesson Lesson slides for teachers |
19 | 18 | 8 | 18 | 6 | 18 | 5 | 3 | 0 | 2 | 3 | 2 |
Job weights: transcription
| Category / Job | Qual | Acc | Spd | Ease | Feat | Value | Rel | Cust | Rights | Sup | Integ | Priv |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
stt.meeting_notes Meeting notes & action items |
13 | 21 | 5 | 6 | 11 | 8 | 5 | 1 | 0 | 2 | 13 | 15 |
stt.interview_multispeaker Interviews & podcasts with multiple speakers |
8 | 34 | 6 | 6 | 9 | 10 | 5 | 2 | 0 | 2 | 2 | 15 |
stt.noisy_field_audio Noisy / field / phone audio |
8 | 38 | 8 | 3 | 3 | 10 | 7 | 2 | 0 | 2 | 2 | 15 |
stt.hebrew_multilingual Hebrew & multilingual / code-switching audio |
8 | 36 | 8 | 6 | 4 | 10 | 5 | 2 | 0 | 2 | 2 | 15 |
stt.api_batch_cost High-volume API transcription on a budget |
4 | 30 | 11 | 2 | 2 | 19 | 8 | 2 | 0 | 2 | 4 | 15 |
Relationship to the v1 weights in research/mvp-decision.md §4. That document expresses weights as plain-language criteria (e.g. Meeting notes: transcript accuracy 25 · summary accuracy 25 · action items 15 · privacy/data 15 · value 10 · integrations 10). The lab weights above are the operational version. Transcript, summary and action-item accuracy map to accuracy, value maps to price_value, and privacy/data maps to privacy_data.
How privacy_data is weighted (CEO decision, v0.2). privacy_data is 15% for every meeting-notes/transcription job and for voice.voice_clone. Elsewhere it is small: voice 4%, chatbots 3%, coding and presentations 2%, video and image 1%, music 0%. The other eleven weights are scaled by (1 − p), so every row still sums to 100%. Rounding remainders go to the largest weight. The source of truth is PRIVACY_DEFAULT / PRIVACY_JOB in lab/build_jobs.py.
5. From outputs to scores
- Task score. Each task has a measured part (0–1 from automatic checks → 1 + 9×value) and/or a judged part (the median across raters of each criterion, then the mean across criteria). Repetitions: the median of 3 (§
_common.md). Refusals score 1 unless refusing is the correct behaviour. - Dimension score. This is the mean of the task-level contributions mapped to that dimension (mapping tables in each category file). Speed, price_value and reliability use the category anchors.
- Job lab score (0–100) = 10 × Σ(wᵈ × dimᵈ) / Σ(wᵈ), summed over the dimensions actually measured. Weight coverage = Σ(wᵈ measured) and is always published. If coverage is below 80%, or any job task is missing, the status is
PARTIALLY_TESTED. - Ranking inside a job is by lab score. Ties within the confidence band are shown as ties, using the same rank and a "statistically indistinguishable" note when the bootstrap 90% intervals overlap.
- All of this is implemented in
lab/harness/harness.py score(API tasks). UI tasks follow the same formulas from manual run directories. Scores are written tojob_scoresby code only; there is no hand entry.
6. Combining lab and user scores
- The lab score and the community score (verified reviews; 1–10 → 0–100) are always displayed separately.
- The optional Decision Score shown on job leaderboards = Lab × (1 − wᵤ) + Community × wᵤ, where wᵤ = min(0.20, 0.20 × n_verified / 50) and wᵤ = 0 if n_verified < 10. User input can therefore move a score by at most 20%, and only once there is real volume.
- Only reviews with
verified=1,moderation_status='approved'andis_affiliated=0count. Reviews are weighted to the job: a review whoseuse_casemaps to the job counts fully, other reviews count 0.25. Spam scoring and rules are indata/schema.md. - MVP: there are no public reviews (research decision), so wᵤ = 0 everywhere.
7. Cost per output
Formulas are in _common.md: API cost = Σ units × unit price; credit plans = price ÷ credits × consumed; flat plans = price ÷ fair-use cap; and the effective cost per usable output, which is cost per attempt ÷ (share of attempts rated ≥ 7), with attempts capped at 4. Prices come from same-day official-page snapshots (data/monitor/). Promotions are excluded. price_value anchors (API LLMs, $ per 1,000 suite items): < $1 → 10 · 1–5 → 8 · 5–20 → 6 · 20–60 → 4 · > 60 → 2. Media categories use per-output anchors defined in each category file.
8. Confidence levels (published next to every score)
| Level | All of these must hold |
|---|---|
| high | every task of the job run · weight coverage ≥ 80% · ≥ 3 repetitions · judged criteria rated by ≥ 2 (target 3) blind human raters · Krippendorff's α ≥ 0.667 on every judged criterion · tested within the freshness window |
| medium | ≥ all-but-one task · coverage ≥ 60% · ≥ 2 repetitions · if judged, ≥ 2 raters (α may be lower, and is shown) |
| low | anything less: a single run, partial tasks, or one rater |
| none | not tested (NOT_YET_TESTED) |
LLM-judge ratings may be stored (evaluator_type='llm_judge'), but they are labelled and never count as raters for confidence or for the judged score. |
9. Blind multi-evaluator rating
- Outputs are exported into blind packs (
harness.py export-blind). Each pack has random IDs, a randomised order per item set and no tool names, file metadata or watermark hints; videos are re-encoded and visible vendor watermarks are cropped or covered where possible. The key is stored separately (_blind_keys/) and only the lab lead can access it. - ≥ 3 raters per judged criterion (minimum 2 for publication at medium confidence). For each category, at least one rater has domain expertise: a producer/musician, a designer, a voice/audio professional, or a native speaker for language tasks.
- Calibration happens before each round: raters score 5 anchor examples and must land within ±1 of the reference on 4 of the 5.
- Agreement is measured with Krippendorff's α (interval) per criterion. If α < 0.667, the criterion is reviewed (rubric clarified, re-rated) before its scores are published, or it is published at low confidence.
- Ratings are imported by
harness.py import-ratings, which enforces a 1–10 scale, rater ids and a written rationale for every score of 1–2 or 9–10. - Raters must not hold affiliate links or have any financial relationship with the vendors they rate. They declare this per round.
10. Conflict-of-interest policy
- No paid placement, ever. Rankings, "best for" badges and inclusion in tests cannot be bought. No sponsored tests.
- Affiliate separation in code. Affiliate status is stored only in
affiliate_programs, and is attached for disclosure after scores are computed. Scoring code,v_public_rankingsandtools.jsonnever contain it. Tool selection for testing ignores affiliate status (research selection rule). - Disclosure sits on every page with an outbound link, along with the statement "affiliate status is never a scoring input".
- Vendor contact. Vendors may report factual errors. These are corrected with a changelog entry. Vendors cannot see results before publication and cannot veto them. Free or press accounts are disclosed on the tool page; where possible we test on paid accounts bought anonymously, at retail.
- Staff. Anyone with equity, employment or an advisory role at a vendor is excluded from rating that category.
- Corrections. Errors are fixed publicly, with date and reason, and the old value is kept in history.
11. Freshness and decay
- Re-test windows (
v_staleness): chatbots, coding and video every 60 days; image, music and voice every 90 days; presentations and transcription every 120 days. Official-page facts are re-checked daily bycheck_pricing.py, and a fact is stale after 30 days without a check. - Triggers for an early re-test: a major version release (e.g. Gamma 5 on 2026-10-06 → re-test Gamma before its first publication), a model swap shown in the UI or API, a price or plan change on the tested tier detected by the monitor, or a credible reported regression.
- Display decay. Scores past their window are marked
STALEand shown with a "results may be outdated" note, and their confidence drops one level. After 2× the window they are removed from rankings (still visible on the tool page as history). Scores are never numerically extrapolated. - Versioning. Each score row stores
methodology_version. A methodology change that would move scores by more than 5 points triggers a re-score from archived outputs where possible, and otherwise a STALE mark.
12. Evidence archiving
- Every run has a directory
data/test_runs/<run_id>/containingmanifest.json(sha256 of every file, suite hash, versions, environment),results.jsonl, raw outputs and ratings. UI runs add screen recordings. - Official pages behind every fact are archived as text plus gzipped HTML with sha256 (
data/monitor/snapshots/) and linked fromfield_sources.snapshot_path. - Evidence is published for every score: outputs in the gallery and prompts in the benchmark files. Exceptions are private test repos and consented voice samples, which are kept private and described instead.
- Retention is indefinite for published scores. Storage: Git for small files, object storage (R2) for media, with the hashes kept in Git.
13. Known limitations (v0.2)
- Automated adapters exist only for LLM and image APIs. Video, avatar, music, voice, presentation and meeting-notes apps are tested manually under a screen-recorded protocol.
- Test assets (original articles, contract, recordings, repos) still have to be produced before the first run; see each benchmark file.
- Consumer pricing for higgsfield, hailuo and kling, and current Hume pricing, could not be read from official pages (re-tried on 2026-10-08 via help centres, developer docs, app-store listings and Wayback; see
lab/freshness-process.md§6). Kling and Hailuo carry official API prices only; Genspark prices now come from its official help centre. Anything missing is marked unknown or low-confidence and is not filled from third parties. - privacy_data facts come from vendor statements on official pages. We can verify deletion and opt-out in our own test accounts, but we cannot audit vendors' internal practices.
Benchmark suites
- Shared benchmark conventions (apply to every category file)
- Benchmark: Chatbots & LLMs (v0.1)
- Benchmark: Coding AI (v0.1)
- Benchmark: Image generation & editing (v0.1)
- Benchmark: Music generation (v0.1)
- Benchmark: AI presentations (v0.1)
- Benchmark: AI meeting notes & transcription (v0.2)
- Benchmark: Video generation (v0.1)
- Benchmark: Voice / TTS & voice cloning (v0.2)
Scoring weights by job
Different jobs value different things. These weights are published in advance.
Cinematic scene with a consistent character
| Dimension | Weight for this job |
|---|---|
| Quality | 38% |
| Accuracy | 17% |
| Features | 10% |
| Customization | 9% |
| Price / value | 6% |
| Reliability | 6% |
| Commercial rights | 6% |
| Speed | 4% |
| Ease of use | 3% |
| Support | 1% |
| Privacy & data handling | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
15-second product ad for Instagram/TikTok
| Dimension | Weight for this job |
|---|---|
| Quality | 22% |
| Accuracy | 15% |
| Speed | 12% |
| Price / value | 12% |
| Ease of use | 11% |
| Features | 10% |
| Commercial rights | 10% |
| Reliability | 4% |
| Customization | 2% |
| Support | 1% |
| Integrations | 1% |
| Privacy & data handling | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Talking-avatar explainer / training video
| Dimension | Weight for this job |
|---|---|
| Accuracy | 20% |
| Quality | 20% |
| Features | 14% |
| Ease of use | 12% |
| Price / value | 8% |
| Speed | 7% |
| Reliability | 6% |
| Integrations | 5% |
| Commercial rights | 4% |
| Customization | 2% |
| Support | 2% |
| Privacy & data handling | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Animate a product photo (image-to-video)
| Dimension | Weight for this job |
|---|---|
| Quality | 32% |
| Accuracy | 21% |
| Price / value | 10% |
| Speed | 7% |
| Ease of use | 7% |
| Features | 7% |
| Reliability | 6% |
| Commercial rights | 6% |
| Customization | 3% |
| Support | 1% |
| Privacy & data handling | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
value video AI for solo creators
| Dimension | Weight for this job |
|---|---|
| Price / value | 25% |
| Quality | 22% |
| Accuracy | 15% |
| Ease of use | 10% |
| Features | 8% |
| Speed | 7% |
| Commercial rights | 6% |
| Reliability | 4% |
| Customization | 2% |
| Support | 1% |
| Privacy & data handling | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Photoreal product & lifestyle shots
| Dimension | Weight for this job |
|---|---|
| Quality | 33% |
| Accuracy | 20% |
| Price / value | 10% |
| Commercial rights | 10% |
| Features | 8% |
| Reliability | 5% |
| Speed | 4% |
| Ease of use | 4% |
| Customization | 4% |
| Support | 2% |
| Privacy & data handling | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Posters, ads & images with correct text
| Dimension | Weight for this job |
|---|---|
| Accuracy | 30% |
| Quality | 23% |
| Price / value | 10% |
| Features | 8% |
| Ease of use | 7% |
| Commercial rights | 7% |
| Speed | 4% |
| Customization | 4% |
| Reliability | 3% |
| Support | 2% |
| Integrations | 2% |
| Privacy & data handling | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Consistent character / brand style across images
| Dimension | Weight for this job |
|---|---|
| Quality | 24% |
| Accuracy | 22% |
| Customization | 15% |
| Price / value | 10% |
| Features | 8% |
| Commercial rights | 7% |
| Reliability | 5% |
| Ease of use | 4% |
| Speed | 3% |
| Support | 2% |
| Privacy & data handling | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
AI photo editing (inpaint, remove, extend)
| Dimension | Weight for this job |
|---|---|
| Accuracy | 24% |
| Quality | 20% |
| Features | 14% |
| Price / value | 8% |
| Ease of use | 7% |
| Speed | 6% |
| Reliability | 5% |
| Customization | 5% |
| Integrations | 5% |
| Commercial rights | 4% |
| Support | 2% |
| Privacy & data handling | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Vector graphics, icons & logo exploration
| Dimension | Weight for this job |
|---|---|
| Quality | 23% |
| Accuracy | 18% |
| Features | 16% |
| Customization | 10% |
| Price / value | 9% |
| Ease of use | 7% |
| Commercial rights | 7% |
| Speed | 3% |
| Reliability | 3% |
| Support | 2% |
| Integrations | 2% |
| Privacy & data handling | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Long-form narration (audiobook, e-learning)
| Dimension | Weight for this job |
|---|---|
| Quality | 37% |
| Accuracy | 14% |
| Price / value | 10% |
| Reliability | 9% |
| Features | 7% |
| Ease of use | 6% |
| Commercial rights | 6% |
| Privacy & data handling | 4% |
| Customization | 4% |
| Speed | 3% |
| Support | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Real-time voice agent (low latency)
| Dimension | Weight for this job |
|---|---|
| Speed | 24% |
| Quality | 17% |
| Accuracy | 14% |
| Price / value | 10% |
| Reliability | 10% |
| Features | 7% |
| Integrations | 5% |
| Privacy & data handling | 4% |
| Ease of use | 3% |
| Commercial rights | 3% |
| Customization | 2% |
| Support | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Cloning a voice (with consent)
| Dimension | Weight for this job |
|---|---|
| Quality | 26% |
| Accuracy | 21% |
| Privacy & data handling | 15% |
| Customization | 8% |
| Price / value | 6% |
| Ease of use | 5% |
| Reliability | 5% |
| Commercial rights | 5% |
| Features | 4% |
| Speed | 3% |
| Support | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Multilingual voiceover & dubbing
| Dimension | Weight for this job |
|---|---|
| Quality | 26% |
| Accuracy | 22% |
| Features | 11% |
| Price / value | 10% |
| Ease of use | 6% |
| Reliability | 6% |
| Commercial rights | 6% |
| Privacy & data handling | 4% |
| Speed | 4% |
| Customization | 3% |
| Support | 2% |
| Integrations | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Short ad / social voiceover
| Dimension | Weight for this job |
|---|---|
| Quality | 29% |
| Price / value | 12% |
| Commercial rights | 12% |
| Accuracy | 11% |
| Ease of use | 10% |
| Features | 7% |
| Customization | 6% |
| Privacy & data handling | 4% |
| Speed | 4% |
| Reliability | 4% |
| Support | 2% |
| Integrations | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Investor pitch deck from a brief
| Dimension | Weight for this job |
|---|---|
| Quality | 32% |
| Accuracy | 17% |
| Ease of use | 14% |
| Price / value | 9% |
| Features | 8% |
| Customization | 7% |
| Speed | 5% |
| Reliability | 3% |
| Support | 2% |
| Integrations | 2% |
| Privacy & data handling | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Turn a long report into a deck
| Dimension | Weight for this job |
|---|---|
| Accuracy | 27% |
| Quality | 20% |
| Ease of use | 10% |
| Price / value | 10% |
| Speed | 8% |
| Features | 8% |
| Integrations | 6% |
| Reliability | 5% |
| Customization | 4% |
| Support | 2% |
| Privacy & data handling | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
On-brand slides inside PowerPoint / Google Slides
| Dimension | Weight for this job |
|---|---|
| Quality | 19% |
| Integrations | 16% |
| Accuracy | 15% |
| Customization | 13% |
| Ease of use | 10% |
| Features | 8% |
| Price / value | 7% |
| Speed | 5% |
| Reliability | 5% |
| Support | 2% |
| Privacy & data handling | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Lesson slides for teachers
| Dimension | Weight for this job |
|---|---|
| Quality | 19% |
| Accuracy | 18% |
| Ease of use | 18% |
| Price / value | 18% |
| Speed | 8% |
| Features | 6% |
| Reliability | 5% |
| Customization | 3% |
| Integrations | 3% |
| Support | 2% |
| Privacy & data handling | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Meeting notes & action items
| Dimension | Weight for this job |
|---|---|
| Accuracy | 21% |
| Privacy & data handling | 15% |
| Quality | 13% |
| Integrations | 13% |
| Features | 11% |
| Price / value | 9% |
| Ease of use | 6% |
| Speed | 5% |
| Reliability | 5% |
| Support | 2% |
| Customization | 1% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Interviews & podcasts with multiple speakers
| Dimension | Weight for this job |
|---|---|
| Accuracy | 34% |
| Privacy & data handling | 15% |
| Price / value | 10% |
| Features | 9% |
| Quality | 9% |
| Speed | 6% |
| Ease of use | 6% |
| Reliability | 5% |
| Customization | 3% |
| Support | 2% |
| Integrations | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Noisy / field / phone audio
| Dimension | Weight for this job |
|---|---|
| Accuracy | 38% |
| Privacy & data handling | 15% |
| Price / value | 10% |
| Quality | 9% |
| Speed | 9% |
| Reliability | 7% |
| Ease of use | 3% |
| Features | 3% |
| Customization | 3% |
| Support | 2% |
| Integrations | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
Hebrew & multilingual / code-switching audio
| Dimension | Weight for this job |
|---|---|
| Accuracy | 37% |
| Privacy & data handling | 15% |
| Price / value | 10% |
| Quality | 9% |
| Speed | 9% |
| Ease of use | 6% |
| Reliability | 5% |
| Features | 4% |
| Customization | 3% |
| Support | 2% |
| Integrations | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.
High-volume API transcription on a budget
| Dimension | Weight for this job |
|---|---|
| Accuracy | 30% |
| Price / value | 19% |
| Privacy & data handling | 15% |
| Speed | 11% |
| Reliability | 9% |
| Quality | 4% |
| Integrations | 4% |
| Features | 3% |
| Customization | 3% |
| Ease of use | 2% |
| Support | 2% |
Methodology v0.2. Weights are job-specific and published before testing. How scoring works.