Live index · 49 models · Updated Oct 2026

Trust & provenance

Methodology

Last updated 2026-10-07 · Full source audit: docs/audit-2026-10-07.md

Inclusion standard

A model is listed only if it meets at least one bar: top-15 weekly inference volume on a public ranking, primary flagship of a tracked lab, or a genuinely frontier capability (state-of-the-art benchmark, new modality). Obscure checkpoints, minor patch bumps, rumors, and community fine-tunes are excluded. See also editorial policy.

What counts as a primary source

In order of strength: the lab's announcement post, official API documentation or pricing page, model card, technical paper, and the lab's official organization repository (e.g. Hugging Face). An official developer console counts as supporting evidence. News coverage is never the sole source for a factual field, and values are never inferred when a source is silent.

Verification levels

Verified - every core field (release date, context, output limit, parameters, license, pricing, modalities, benchmarks where published, flagship/latest status) was read and confirmed against live primary sources.

Partially verified - one or more live official sources are cited, but full per-field rechecking is still pending. Figures are transcribed from the cited pages.

Unverified - no working official source. Figures must be treated as unconfirmed; such records can never carry flagship or latest labels.

Retired - superseded and kept for reference only, with a pointer to its replacement.

What “flagship”, “latest”, and leaderboard places mean

Flagship = the lab's primary general-purpose model (exactly one per lab). Latest checkpoint = the lab's newest shipped release, which may be a specialized model rather than the flagship. Both labels require fresh (≤ 90 days), live evidence, enforced automatically - stale labels fail validation. Leaderboard sections are an editorial snapshot, not objective fact: leaders hold the highest published score in that comparison as of the stated evaluation date. We say “highest published score in this comparison”, never “best model”.

Benchmarks: lab-published vs independent

Benchmark figures on model pages are lab-published scores as reported by the vendor, kept separate per harness and version (DeepSWE v1.1 is not SWE-bench Verified). They are not independently reproduced by this registry. Independent evaluations (e.g. Artificial Analysis indices) are cited as such where used. Figures without a citable source are shown with an explicit “no citation yet” marker instead of a borrowed number.

Downgrade, correction, retirement, removal

A record is downgraded (e.g. flagship → previous generation, or a label removed) when its evidence lapses; corrected when a primary source contradicts a field; retired when superseded (with a pointer); removed only for duplicates, hoaxes, or records that never had a verifiable official source. Every change is recorded in the model's changelog and the public changelog.

Maintainers and corrections

ModelRegistry is a community project maintained via its public GitHub repository. Verification is source review by maintainers - not independent laboratory testing, and we do not claim otherwise. To report an error, use the “Report an error in this record” link on any model page, which opens a correction issue requiring an official source. High-severity factual errors (wrong price, wrong flagship) are fixed within days; routine re-verification follows the schedule below.

Maintenance schedule

Every 30 days - automated liveness check of all cited sources; dead links re-sourced or records downgraded.

Every 60 days - pricing and context figures re-checked against official pricing/docs pages for all current flagships and latest checkpoints.

Every 90 days - full freshness rotation: any flagship/latest label older than 90 days fails validation until re-verified; leaderboard comparisons re-evaluated and re-dated; a new dated audit note is published.