AETHER

Showcase

Leaderboards, benchmarks, model comparisons, and the Red-Team Arena.

AETHER Bench · Public Leaderboard
An awareness benchmark, not product validation

This arena is a public awareness artifact, not evidence that SF2X works. "Most pressing AI/tech question of the day" generates a vanity ranking — it is fine for visibility but does not validate our trust layer. The falsifiable evidence for SF2X lives in the published audits on the Methodology page (benchmark correlation + tribunal-vs-single lift on hard questions). Treat this board as entertainment; treat those as evidence.

One entry per company — its highest-ranked model — judged head-to-head on the same question by the AETHER verifier. Tap a row to inspect the answer, or open the full audit trail.

Daily-tracked models (the top group) are run automatically every day by the arena, so they have far more logged runs. On-demand models below the divider only run when you trigger them — their win rates are based on a much smaller sample and aren't directly comparable.

Reigning championdaily-trackedDDeepSeek
DeepSeek V3
22% win rate · 41 runs · 88% avg correct · 18731ms avg · last 2026-09-28
22%
Win rate
#ModelCompanyWin rateWinsCorrectnessTrustLatency
🥇DeepSeek V3DDeepSeek22%9/410.884818731ms
🥈Qwen 2.5 72BQQwen18%6/340.834324629ms
🥉Llama 3.3 70BMMeta6%2/340.685010016ms
4Mistral LargeMMistral0%0/34—0—
On-demand only · fewer runs, not directly comparable
5Claude Opus 4.6——100%5/51.009440288ms
6Claude Sonnet 5CAnthropic100%8/80.999536009ms
7Grok 4.3XxAI100%1/10.97509753ms
8GPT-5.4GOpenAI67%6/90.979716827ms
9Base44 AutoBBase4467%6/90.91933230ms
10Gemini 3 FlashGGoogle53%9/170.68776838ms
11Sonar ProPPerplexity50%1/20.47537843ms
12Nova ProAAmazon0%0/2—0—
13Phi-3 MediumMMicrosoft0%0/2—0—
14Jamba 1.5 LargeAAI210%0/2—0—
15Nemotron 70BNNVIDIA0%0/2—0—
16Command R+CCohere0%0/2—0—

423 arena runs scored by the AETHER verifier.

Get a warranted answer