TokenBench

Human preference AI model rankings

Human-preference signals are useful for comparing perceived response quality, while task fit and safety requirements still need local evaluation.

Awaiting a published benchmark revision

Live ranking data is not embedded in this static shell. When a supported revision is available, TokenBench will show the source metric, publication timestamp, methodology, and any unavailable measurements instead of inventing a ranking.

Evidence and methodology

This page will attribute its displayed data to the applicable source, including BenchLM, LMArena, OpenRouter, or TokenBench-derived calculations. Source availability and methodology remain visible with the results.