Roseram NewsAll stories

Chinese & International · OpenCompass

DeepSeek forced everyone to re-read OpenCompass.

Cost, reasoning traces and open-weight releases turned DeepSeek into a global story. The scoreboard is real; the geopolitics is louder.

Tools: DeepSeek · Qwen · Llama

What changed

DeepSeek’s reasoning-model cycle pushed Western desks to check OpenCompass and related leaderboards instead of assuming US-lab defaults. Coverage on Medium’s OpenCompass distillations and daily.dev briefs amplified the scoreboard—sometimes past what the methods section supports.

Roseram reports DeepSeek results with the benchmark suite name, whether tools were allowed, and whether the checkpoint is open-weight or API-only.

Llama remains the foil

Llama still anchors much of the open-weight conversation outside China. Comparing DeepSeek and Llama without matching parameter class, license and eval harness produces false binaries.

Read the underlying record.

  1. 01OpenCompass rankings coverage ↗
  2. 02Ken Huang — test-time compute ↗
  3. 03Stanford AI Index ↗