Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
benchmarks
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Ranking Language Models by How Well They Spot Liars
Seth Wheeler
Seth Wheeler
Seth Wheeler
Follow
Aug 18
Ranking Language Models by How Well They Spot Liars
#
llm
#
measurement
#
benchmarks
#
statistics
Comments
Add Comment
9 min read
The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It
Charles
Charles
Charles
Follow
Aug 18
The Benchmarkpocalypse: Why AI Benchmarks Are Broken — and What Dan Luu Says We Should Do About It
#
ai
#
machinelearning
#
benchmarks
#
programming
1
 reaction
Comments
Add Comment
4 min read
Labs are ditching factual knowledge for reasoning speed
Peremptory
Peremptory
Peremptory
Follow
Aug 18
Labs are ditching factual knowledge for reasoning speed
#
benchmarks
#
modelreleases
#
aistrategy
#
research
Comments
Add Comment
2 min read
Why AI Benchmarks Mean Less Than You Think
The AI Downside
The AI Downside
The AI Downside
Follow
Aug 15
Why AI Benchmarks Mean Less Than You Think
#
benchmarks
#
llms
#
evaluation
#
hype
Comments
Add Comment
6 min read
An AI Capture-the-Flag Tournament: What the Scoreboard Counted
Seth Wheeler
Seth Wheeler
Seth Wheeler
Follow
Aug 15
An AI Capture-the-Flag Tournament: What the Scoreboard Counted
#
llm
#
measurement
#
security
#
benchmarks
Comments
Add Comment
6 min read
Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.
Tommy Leonhardsen
Tommy Leonhardsen
Tommy Leonhardsen
Follow
Aug 11
Twelve LLMs Played Werewolf. The Real Wolf Was the Thinking Knob.
#
llm
#
benchmarks
#
reasoning
#
latency
1
 reaction
Comments
Add Comment
8 min read
Meta's Muse Code Clears 59% on Deep Software Engineering
Peremptory
Peremptory
Peremptory
Follow
Aug 7
Meta's Muse Code Clears 59% on Deep Software Engineering
#
modelrelease
#
developertools
#
benchmarks
#
codingmodels
Comments
Add Comment
2 min read
Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)
Lola Lin
Lola Lin
Lola Lin
Follow
Jul 28
Kimi K3 for Coding: Real-World Performance Tests and Benchmarks (2026)
#
kimik3
#
coding
#
benchmarks
#
softwareengineering
Comments
Add Comment
9 min read
SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 17
SynthDocBench: A New Benchmark for Long-Context Visual Document Understanding Reveals VLM Weaknesses
#
visionlanguagemodels
#
vlm
#
documentunderstanding
#
benchmarks
Comments
Add Comment
4 min read
UniClawBench: A New Benchmark for Proactive AI Agents in Real-World Scenarios
Pneumetron
Pneumetron
Pneumetron
Follow