Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evals
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Evals are the unit tests of prompts
INTFRAME
INTFRAME
INTFRAME
Follow
Aug 15
Evals are the unit tests of prompts
#
ai
#
evals
Comments
Add Comment
2 min read
Predicting agent failure before you ship it
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 14
Predicting agent failure before you ship it
#
failuremodes
#
testing
#
evals
#
reliability
Comments
Add Comment
6 min read
Loop drift: how agents convince themselves they're making progress
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 13
Loop drift: how agents convince themselves they're making progress
#
failuremodes
#
evals
#
postmortem
#
loops
Comments
Add Comment
7 min read
One pass of my eval bills $9.14 on the API and $0 through the CLI
Dylan Merigaud
Dylan Merigaud
Dylan Merigaud
Follow
Aug 12
One pass of my eval bills $9.14 on the API and $0 through the CLI
#
ai
#
evals
#
claude
#
tooling
Comments
Add Comment
2 min read
Your LLM-as-judge is lying to you
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 12
Your LLM-as-judge is lying to you
#
evals
#
llmasjudge
#
testing
#
bias
Comments
Add Comment
8 min read
Evaluating your evals: how to know the LLM judge is right
Walker Miller
Walker Miller
Walker Miller
Follow
Aug 10
Evaluating your evals: how to know the LLM judge is right
#
evals
#
llmasjudge
#
testing
#
metrics
Comments
Add Comment
5 min read
AI per developer: cosa accelera davvero (e cosa ti fa perdere tempo)
frontendfacile.it
frontendfacile.it
frontendfacile.it
Follow
Aug 7
AI per developer: cosa accelera davvero (e cosa ti fa perdere tempo)
#
workflowaicoding
#
agentillm
#
evals
#
reliabilitytesting
Comments
Add Comment
4 min read
I have been Vibecoding Evals (works better than I thought)
juan pablo hernández
juan pablo hernández
juan pablo hernández
Follow
Aug 3
I have been Vibecoding Evals (works better than I thought)
#
ai
#
llm
#
python
#
evals
Comments
Add Comment
3 min read
Rewriting prose until the tests pass: everything passed, but the check that mattered never ran once
matsumotory
matsumotory
matsumotory
Follow
Aug 4
Rewriting prose until the tests pass: everything passed, but the check that mattered never ran once
#
testing
#
aiwriting
#
evals
#
promptengineering
Comments
Add Comment
6 min read
Rewriting research prose until the tests pass
matsumotory
matsumotory
matsumotory
Follow
Aug 4
Rewriting research prose until the tests pass
#
testing
#
research
#
aiwriting
#
evals
Comments
Add Comment
5 min read
The 12-Prompt Eval I Run Before I Trust Any Model Upgrade
Agnel Nieves
Agnel Nieves
Agnel Nieves
Follow
for
Promptway
Jul 29
The 12-Prompt Eval I Run Before I Trust Any Model Upgrade
#
prompting
#
evals
#
modelmigration
#
claude
Comments
Add Comment
3 min read
How to Build AI Evals for Tool-Calling Agents
Dhanush Reddy
Dhanush Reddy
Dhanush Reddy
Follow
Aug 8
How to Build AI Evals for Tool-Calling Agents
#
ai
#
evals
#
aievals
#
agents
1
 reaction
Comments
2
 comments
17 min read
Do not choose an AI model from a leaderboard alone
Edward Li
Edward Li
Edward Li
Follow
Jul 8
Do not choose an AI model from a leaderboard alone
#
ai
#
api
#
llm
#
evals
Comments
Add Comment
3 min read
# A 94% pass rate hid a PII leak in 6 test cases
Ethan Walker