Dan Shipper is the CEO of Every. Mike Taylor, its head of evals, fed Jev 37 documents with 21 questions each and got 777 judgments back in 0.7 seconds for about a quarter of a cent. Every's second test is the one independent number so far: on 12 passages with planted writing defects Jev caught six of seven where Fable 5.1 caught all seven, at 0.35 seconds a passage against 8.83, which is 25 times faster rather than the 200 in the launch claim.
we almost never test new foundation models but we've been testing this for ~a week @every and it's pretty wild.
the kind of things that will be obviously indispensible in 6-12 months
it doesn't produce words as output, it produces probabilities. so it can efficiently act as a judge in cases where you'd need a Fable-level model—but in our testing was 25x faster and 600x lower priced
excellent vibe check by @hammer_mt on @every: