Skip to main content
You can combine vision input with structured outputs to extract typed data from an image. Pass an image_url content block and a response_format with a JSON schema; the model returns JSON that conforms to the schema. For example, you could extract a project name and a column count from a screenshot of a Trello board:
Example output:
JSON

OCR: extract text from documents

Optical character recognition (OCR) is the same pattern applied to documents. Vision models read the text in an image while understanding its context and structure, which makes them useful for processing receipts, invoices, and other structured documents. See llamaOCR.com for a working example. To extract everything a document contains as plain text, send the image without a schema:
The model returns the receipt’s contents as formatted text:
Text

Extract structured data from a receipt

For receipt processing (as seen on usebillsplit.com), combine OCR with a schema to get back typed JSON instead of free text:
The response contains only the fields the schema asks for:
JSON

Best practices

  1. Structured data definition: Define clear schemas for your expected output, so you can validate and process the extracted data.
  2. Model selection: Choose the appropriate model based on your use case. Try Together’s vision models to match your workload.
  3. Error handling: Always implement robust error handling for cases where the OCR might fail or return unexpected results.
  4. Validation: Implement validation for the extracted data to ensure accuracy and completeness.
For the full structured-outputs reference, see Structured outputs.