What does the Nutrient Data Extraction API do?
The Nutrient Data Extraction API is a service that parses PDFs, scans, images, and Office files into structured JSON or Markdown. It extracts structured data with coordinates, confidence scores, and page context for use in agents, RAG, search, automation, and human review.
What problem does this API solve for developers?
It solves the problem of extracting structured data from complex, real-world documents like PDFs, scans, and Office files into a predictable format for downstream systems. This enables building deterministic workflows for AI and automation, moving from document parsing to extraction and structured output.
Who is this API designed for?
This API is designed for developers and teams building document workflows, automation, and AI applications. It is used by enterprises, governments, and AI-native teams, with customers including Lufthansa, Disney, Autodesk, UBS, Dropbox, and IBM.
What are the key features and capabilities of the API?
Key capabilities include detecting every document element (tables, forms, formulas, images, handwriting), preserving page context and reading order, choosing between speed, cost, or depth processing modes, and handling real-world files. It also allows mapping data to a custom schema and extracting complex tables with rows, columns, and spans preserved.
What is the output format of the extracted data?
The API returns extracted data as spatial JSON or Markdown. The output includes bounding boxes, match labels, confidence scores, page position, and source evidence for review, making it ready to map into databases, ERPs, CRMs, and other downstream systems.
What file types can the Nutrient Data Extraction API process?
The API can process PDFs, scans, images, Office files, and multilingual documents. It handles real-world files through a single API endpoint.
How is the Nutrient Data Extraction API priced?
The API offers a free starting plan with 5,000 monthly credits. Further pricing details for scales beyond the free tier are not specified in the provided evidence.
What product category does this API belong to?
This API is categorized as a Developer & AI Platform. It is an API service for extracting structured data from documents, targeting developers building document workflows and automation.