Model rankings
A curated summary of well-regarded models by task and size, drawn from external benchmarks. We do not run our own benchmarks; every entry is attributed and dated so you can judge it yourself.
General chat
General-purpose assistants and instruction following.
| Size tier | Model | Draws on | Source | As of |
|---|---|---|---|---|
| Large | Llama 3.3 70B Instruct Meta A releasing-organisation figure. Treat as a claim pending independent confirmation. | MMLU 86.0 | Meta model card | December 2024 |
Coding
Code generation, completion, and understanding.
No published assessment yet. We will add one here once we have a benchmark and source we can stand behind.
Reasoning
Maths, logic, and multi-step problem solving.
No published assessment yet. We will add one here once we have a benchmark and source we can stand behind.
Embedding
Text embeddings for retrieval and RAG.
No published assessment yet. We will add one here once we have a benchmark and source we can stand behind.
Vision-language
Understanding images alongside text.
No published assessment yet. We will add one here once we have a benchmark and source we can stand behind.
Image generation
Generating images from text prompts.
No published assessment yet. We will add one here once we have a benchmark and source we can stand behind.
Speech
Speech recognition and synthesis.
No published assessment yet. We will add one here once we have a benchmark and source we can stand behind.
Rankings are curated and human-approved. Our pipeline can gather current benchmark figures for review, but nothing is published here without a source and a date.