Perslis Accessibility
04 / SEE TO SPEAK

At the shops. At the counter. Point, and get the words.

A board can only hold the words someone put on it in advance. A camera removes that limit: send a photo and the tiles become the thing in front of the person — this menu, this shelf, this object — in sentences they can say.

THREE KINDS OF LOOKING

It works out which one you sent.

Send mode: "auto" and it decides between an object, a menu and a scene; name the mode yourself when you already know.

InputWhat comes back
An object — shoes, gloves, foodObject names, body-use associations where relevant, and choices for requests, help, actions and refusal.
A menu photographThe item names and prices it reads off the page, with categories and ordering sentences. Several pages stay in one catalog, and a misread is surfaced rather than hidden.
A sceneVisible objects, optional image rectangles and spatial relations, plus choices relevant to that scene.
/object "/path/to/shoes.jpg"
/menu "/path/to/menu-1.jpg" "/path/to/menu-2.jpg"
/scene "/path/to/room.jpg"
/see "/path/to/photo.jpg"
THE CONVERSATION CONTINUES

Seeing is not a dead end.

After a scan, ordinary partner input keeps using the vision model together with the saved observations: shoes → “I need my shoes” → “Do you need help?” → relevant follow-up choices. No new image is sent, and the runtime does not claim to observe a changed scene. /unsee returns to the text model and keeps the conversation.

WHAT THE DATA PROMISES

Observations, with their limits stated.

01

Boxes are fractions

{x,y,w,h} of image width and height, top-left origin. Model observations only — no depth estimate and no navigation guarantee.

02

Instances stay distinct

Relations expose indices into the object list, so two cups are two cups. An ambiguous group name yields a null index rather than a guess.

03

Prices are display data

Separate from the proposed sentence. Selecting an ordering tile speaks the user's words; it never submits a purchase or payment.

04

Wrong is reported, not hidden

Recognition and OCR can be wrong. Warnings and the original item text are exposed for your review interface. An incomplete result fails explicitly rather than substituting a fabricated menu.

05

Limits are explicit

Up to eight images per scan, 5 MiB each and 12 MiB combined. Sentences are never truncated; replies that break the word limit are reported as excluded.

06

Photos are not kept

Image bytes exist only during the request. Session snapshots keep hashes, type, size and observations — never the image, a path or base64.

Offline and images are different questions

The text-only child model does not gain image recognition by configuration. A local Ollama model can process images on your host, but it uses HTTP, so --offline disables that connection too. A loopback endpoint alone does not prove offline inference.

ContinueModels & privacy

Everything Perslis