Alex Duffy runs Goodstart Labs, which builds learning environments for agents and needs every action graded. On their sample Jev agreed with Fable 5.1 nine times in ten, at about 200 times lower grading cost and under half a second a call. Erik Gafni, named in the post, is one of TypeSafe's three co-founders.
Awesome to see @EGafni & Diogo launch their first model Jev!
Verification is the bottleneck for building great AI learning environments. Jev is a welcome, unique, addition to our toolkit @goodstartlabs
Typed decisions,
probabilities for each option,
free output tokens!
On our sample:
• Jev agreed with Fable 5.1 nine times out of ten
• ~200× lower grading cost
• <0.5 seconds per grading call
That means more eyes on what agents are doing and what we’re teaching them.
More on how we’re using it to build game environments, experts to play them, and ensuring they teach the right things: