Okay so Jev can actually do computer use really well
Without any screenshots, or LLMs and no Pixels leave my mac
I dont even read the Dom elements
A local CoreML model segments every button and UI element on screen.
On-device OCR reads the labels. That text is all Jev gets.
It returns a probability across those elements and tells me the best one to click.
Then it clicks, re-runs detection, and decides again. In a loop until the goal is done.
~90ms per decision. Faster than any LLM computer use I've tried.
Blazing fast computer use, without any latency
@typesafeai is building something really interesting