DEV Community

#benchmarks

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Ranking Language Models by How Well They Spot Liars

Ranking Language Models by How Well They Spot Liars

Comments
9 min read
The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It

The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It

1
Comments
4 min read
Labs are ditching factual knowledge for reasoning speed

Labs are ditching factual knowledge for reasoning speed

Comments
2 min read
Why AI Benchmarks Mean Less Than You Think

Why AI Benchmarks Mean Less Than You Think

Comments
6 min read
An AI Capture-the-Flag Tournament: What the Scoreboard Counted

An AI Capture-the-Flag Tournament: What the Scoreboard Counted

Comments
6 min read
Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.

Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.

1
Comments
8 min read
Meta's Muse Code Clears 59% on Deep Software Engineering

Meta's Muse Code Clears 59% on Deep Software Engineering

Comments
2 min read
Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)

Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)

Comments
9 min read
SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses

SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses

Comments
4 min read
UniClawBench: A New Benchmark for Proactive AI Agents in Real-World Scenarios