July 10, 2026
The Eval Stack: Proving the agents are right instead of claiming It
Most AI research tools point a language model at the web and trust the output. Saarth Shah built Sixtyfour around the opposite instinct: grade everything, ship only what improves the score. Saarth Shah keeps a scoreboard. Every build of Sixtyfour’s r...