DEV Community

Agent Determinism Illusions Series' Articles

Back to zxpmail's Series
I tested the 'deterministic agent loop' claims with four experiments. They all failed — including my own fix.

Formal code failing at the semantic layer

I tested the 'deterministic agent loop' claims with four experiments. They all failed — including my own fix.

10
Comments 10
9 min read
I tested 3 models as AI agent quality inspectors: the stronger the model, the more valid work it rejects

Precision-recall and the false positive mirage

I tested 3 models as AI agent quality inspectors: the stronger the model, the more valid work it rejects

10
Comments 60
5 min read
I designed a Harness to fix my agent's quality problem — then found 6 flaws in my own design

I designed a Harness to fix my agent's quality problem — then found 6 flaws in my own design

3
Comments 5
8 min read
An alternative to LLM quality gates: deterministic routing + sampling

Six experiments on LLM-judge failures

An alternative to LLM quality gates: deterministic routing + sampling

13
Comments 69
11 min read
Six experiments on adversarial verification — and the 75% wall that didn't move

Six experiments on adversarial verification — and the 75% wall that didn't move

12
Comments 44
6 min read
The Red Line Principle: objective stop signals outperform LLM self-judgment in verifiable tasks

The Red Line Principle: objective stop signals outperform LLM self-judgment in verifiable tasks

4
Comments 12
13 min read
Five Comments That Redesigned My LLM Verification Pipeline

Five Comments That Redesigned My LLM Verification Pipeline

5
Comments 44
26 min read
Divergence escalates the wrong population: unanimous misses auto-pass

Divergence escalates the wrong population: unanimous misses auto-pass

3
Comments 28
17 min read
I Fabricated a Claim About LLM Judges. Then I Ran the Apology Experiment.

Rigorous self-correction via 600 test calls

I Fabricated a Claim About LLM Judges. Then I Ran the Apology Experiment.

1
Comments 5
8 min read
D+T2 names who enters; budget names who gets seen

D+T2 names who enters; budget names who gets seen

3
Comments 10
6 min read
Reader-driven revisions: four comments that bit back

Reader-driven revisions: four comments that bit back

3
Comments
5 min read
Round 2: when the reply triggers another revision

Round 2: when the reply triggers another revision

3
Comments 2
4 min read
The Channel Gap: Why Your LLM Judge is Blind in One Eye

Combines text judges with filesystem checks

The Channel Gap: Why Your LLM Judge is Blind in One Eye

20
Comments 19
15 min read
Weng's Harness Ladder Has a Blind Step

Weng's Harness Ladder Has a Blind Step

9
Comments 13
18 min read
The Third Predicate: Argument-Space Verification, Tested

The Third Predicate: Argument-Space Verification, Tested

4
Comments 2
15 min read