Computer use
blink can pick an agent's next UI action as a typed choice. It reads page elements as text, or screenshots on a self-hosted server.
Try it
| Link | What it does |
|---|---|
| Screen click | Picks the next click on a screenshot with blink-mimo-9b. |
| Open the result | Opens a saved catalog run. |
| Filter first | Opens a saved directory run. |
| Pick a topic | Opens a saved lookup run. |
| Already done | Opens a saved completed-task run. |
| Next click | Picks the next action from a text list of page elements. |
Run it yourself
| Link | Setup |
|---|---|
| Screenshot input | Self-host serve.py. Use --vision for blink-mimo-9b or --vision-tower for blink-4b and blink-27b. |
| blink-4b screenshot setup | Use the base model's vision tower. |
| blink-mimo-9b screenshot setup | Use its own vision tower. |
| blink-27b screenshot setup | Use the base model's vision tower. |
| Browser agents | Use jev-ultrafast as the loop and blink as the decision server. Apply the same two-line base-URL change. |
Good to know
- Screen click is a demo and never clicks anything. Presets show saved runs from blink-mimo-9b. Uploads run live.
- The Space API endpoint stays text-only. Screenshots work with self-hosted
serve.py, where image input is off by default. - TypeSafe's hosted Jev is text-only. Image input is a self-hosted blink extension.
- The opt-in
serve_vllm.pyfor blink-4b is text-only.