DEV Community

#statistics

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Ranking Language Models by How Well They Spot Liars

Ranking Language Models by How Well They Spot Liars

Comments
9 min read
A judge that agrees with your humans 92 percent of the time can be at 60 percent where the gate actually decides

A judge that agrees with your humans 92 percent of the time can be at 60 percent where the gate actually decides

1
Comments
8 min read
Multi-Armed Bandit Testing: How It Works and When to Use

Multi-Armed Bandit Testing: How It Works and When to Use

Comments
11 min read
How to Choose a Minimum Detectable Effect (MDE)

How to Choose a Minimum Detectable Effect (MDE)

Comments
9 min read
You Don't Need a Math PhD for Data Science — You Need to Stop Skipping the Boring Step

You Don't Need a Math PhD for Data Science — You Need to Stop Skipping the Boring Step

Comments
4 min read
Your eval monitor fired on four days this week. At your sample size, that was the most likely count

Your eval monitor fired on four days this week. At your sample size, that was the most likely count

Comments
8 min read
Part 6 - STATISTICS

Part 6 - STATISTICS

Comments
8 min read
Your eval monitor fired on four days this week. At your sample size, that was the most likely count

Your eval monitor fired on four days this week. At your sample size, that was the most likely count

1
Comments
8 min read
Upgrading the judge ends one score series and starts another

Upgrading the judge ends one score series and starts another

5
Comments
9 min read
AI told me a method didn't work. It had barely run it.

AI told me a method didn't work. It had barely run it.

Comments
18 min read
Beyond the Mean: An Engineer's Guide to Making Decisions with Statistical Inference

Beyond the Mean: An Engineer's Guide to Making Decisions with Statistical Inference

1
Comments
11 min read
Why "Accuracy" Is the Wrong Metric for Probabilistic Prediction Models

Why "Accuracy" Is the Wrong Metric for Probabilistic Prediction Models

Comments
4 min read
Your A/B test has three goals and they disagree. Now what?

Your A/B test has three goals and they disagree. Now what?

Comments
7 min read
A noisy judge does not just add error bars. It shrinks the effect you are trying to measure.

A noisy judge does not just add error bars. It shrinks the effect you are trying to measure.

1
Comments
11 min read
Part 5 - STATISTICS

Part 5 - STATISTICS

Comments
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.