blink

Computer use

blink can pick an agent's next UI action as a typed choice. It reads page elements as text, or screenshots on a self-hosted server.

Try it

Link What it does
Screen click Picks the next click on a screenshot with blink-mimo-9b.
Open the result Opens a saved catalog run.
Filter first Opens a saved directory run.
Pick a topic Opens a saved lookup run.
Already done Opens a saved completed-task run.
Next click Picks the next action from a text list of page elements.

Run it yourself

Link Setup
Screenshot input Self-host serve.py. Use --vision for blink-mimo-9b or --vision-tower for blink-4b and blink-27b.
blink-4b screenshot setup Use the base model's vision tower.
blink-mimo-9b screenshot setup Use its own vision tower.
blink-27b screenshot setup Use the base model's vision tower.
Browser agents Use jev-ultrafast as the loop and blink as the decision server. Apply the same two-line base-URL change.

Good to know

  • Screen click is a demo and never clicks anything. Presets show saved runs from blink-mimo-9b. Uploads run live.
  • The Space API endpoint stays text-only. Screenshots work with self-hosted serve.py, where image input is off by default.
  • TypeSafe's hosted Jev is text-only. Image input is a self-hosted blink extension.
  • The opt-in serve_vllm.py for blink-4b is text-only.

View as MarkdownView source on GitHub