
The Saw Test: What a Small Model Does When Relief Costs Someone Else
The validator wars: three SSRF/input validators, three bypasses
Reconstructing undisclosed Plex fixes by behavioral diff
Three live proofs from the AI-infra audit round
Tireless search beats cleverness: a lab for auditing self-hosted AI
Reproducing "Fractal basins trap latent reasoning" on a 27M-param recurrent solver