← Back to blog

OCR API for Brazilian documents: complete guide

Learn how to extract data from birth certificates with the DocsOCR API. Complete guide with code examples.

What is document OCR and why does it matter?

OCR (Optical Character Recognition) is the technology that transforms text in images — such as scanned documents, photos, and PDFs — into editable, searchable digital data. For Brazilian companies handling large volumes of documents, OCR is the difference between processing hundreds of documents per day or being stuck with manual data entry.

In Brazil, the digitization of processes has accelerated significantly. Registry offices, fintechs, law firms, and HR companies need to extract data from documents like the birth certificate at scale. The problem? Most generic OCR solutions weren’t designed to handle the particularities of Brazilian documents.

How DocsOCR works

DocsOCR is a REST API specialized in Brazilian documents, built for the official model of the birth certificate.

Three AI engines with fallback

The system uses three AI engines with automatic fallback:

  • Fast: the engine a request can ask for with "engine": "fast"; it is tried first, at its own price
  • Large (Accurate): a standard engine, optimized for precision in field extraction
  • Mini (Balanced): a standard engine, for high volume

Mini and Large are the standard engines, which answer every request by default. When an engine fails on a document, or is unavailable, the system automatically tries the next one, with no manual step. If none can answer, the answer comes back as 201 with success: false and an errorCode, at no charge: send it again later.

Processing flow

  1. Upload: You send the certificate image via API (URL or Base64)
  2. Extraction: The AI engines read the certificate and extract its fields
  3. Response: the data comes back as structured JSON, in the same request

Cost of each engine

An extraction on the standard engines costs 1 credit, with nothing extra when the system switches engines because one of them is unavailable. Speed is your choice: a request can ask for the fast engine with "engine": "fast", and then pays fast’s own price when fast answers. The API’s price list states it at GET /documents/prices. If fast cannot answer, a standard engine does, for 1 credit. And you pay only for extracted data: a refused request, or one that no engine could serve, costs nothing.

Image quality matters. A straight, sharp, well-lit photo of the whole page gives the best results; small or blurry photos can lose accents and digits.

Supported documents

Currently, DocsOCR offers full extraction for Brazilian birth certificates, including:

  • Full name of the registered person
  • Date of birth
  • Place of birth (city and state)
  • Parents’ names
  • Registration and enrollment numbers
  • Issuing registry office name
  • Date of issuance
  • Supplementary data (grandparents, notes)

More document types coming soon.

Getting started with the API

Getting started with DocsOCR is simple:

1. Create your account

Access the admin panel and create a free account. You’ll receive welcome credits to test the API immediately — no credit card required.

2. Generate an API key

In the dashboard, navigate to the API keys section and generate your first key. Keep it in a safe place.

3. Make your first call

Put the key in the DOCSOCR_API_KEY environment variable and run this Python example:

"""Your first call: extract the hosted sample certificate (fictitious data).

Needs the requests package and DOCSOCR_API_KEY in the environment.
"""
import os
import uuid

import requests

response = requests.post(
    "https://api.docsocr.com/api/v1/documents/birth-certificate",
    headers={"Authorization": f"Bearer {os.environ['DOCSOCR_API_KEY']}"},
    json={
        "imageType": "url",
        "imageUrl": "https://docsocr.com/samples/certidao-nascimento-exemplo.jpg",
        "requestId": str(uuid.uuid4()),  # one id per document
    },
    timeout=120,  # above the API's 90 s extraction budget
)
answer = response.json()
if response.status_code != 201 or not answer["success"]:
    raise SystemExit(f"{response.status_code} {answer.get('errorCode', '')} {answer.get('message') or answer.get('error')}")
print(answer["data"]["dados_pessoais"]["nome_completo"])

Or with cURL:

#!/usr/bin/env bash
# Your first call: extract the hosted sample certificate (fictitious data).
# Needs DOCSOCR_API_KEY in the environment; create a key in the panel.
set -euo pipefail

REQUEST_ID="first-call-$(date +%s)-$RANDOM" # one id per document

# The API answers within 90 s
curl -sS --max-time 120 https://api.docsocr.com/api/v1/documents/birth-certificate \
  -H "Authorization: Bearer $DOCSOCR_API_KEY" \
  -H "Content-Type: application/json" \
  --data-binary @- <<EOF
{
  "imageType": "url",
  "imageUrl": "https://docsocr.com/samples/certidao-nascimento-exemplo.jpg",
  "requestId": "$REQUEST_ID"
}
EOF

The response includes all extracted fields in structured JSON format, ready to be integrated into your system.

The examples use a fictitious certificate we host for testing. For your own documents, change the URL or send the image as Base64. Each request carries its own requestId. The same examples, and five more (local file, resizeImage, the fast engine, errors and prices), come in six languages.

Accepted formats

DocsOCR accepts the following file formats:

  • PDF* — digital or scanned documents
  • JPG/JPEG — photos and scans
  • PNG — images with or without transparency
  • WebP — compressed images, common on the web and on Android phones
  • GIF — still images

*Only the first page of a PDF is read.

The maximum file size is 10 MB. You can send the document as a public URL or Base64-encoded content in the API request body.

Our image standard is 1344 to 2048 px on the long side, where extraction reads best. An image outside it is refused at no charge, or, if the request sets resizeImage: true, resized on our side for 1 extra credit, which adds a little time to the answer. For the fastest and cheapest answer, resize before sending: a phone photo at 2048 px on the long side is enough.

If the image cannot be used, or the URL cannot be downloaded, the API answers with HTTP 422, an errorCode that names the cause (for example IMAGE_TOO_LARGE) and a message that explains the cause and the fix: for example, use a public, direct link to the file, or send the image as Base64.

Pricing and free credits

DocsOCR uses a pay-as-you-go credit model:

  • Each extraction that returns data costs 1 credit, or fast’s price when the request asks for the fast engine and fast answers; a refused request, or one that no engine could serve, is free
  • New accounts receive free credits to test with real documents
  • No credit card required to get started
  • Monthly and annual plans available, with progressive volume discounts

The pay-as-you-go model ensures you only pay for what you use, with no high fixed costs or billing surprises.

Security and compliance

Data security is a top priority:

  • LGPD, GDPR, and CCPA: 100% compliant with major data protection regulations
  • Real-time processing: Document images are not stored after extraction
  • TLS encryption: All communication is encrypted in transit
  • API keys per workspace: you can revoke a key at any time in the dashboard

Next steps

Ready to automate document processing in your company?

  1. Create your free account and receive welcome credits
  2. Explore the interactive documentation and the examples in six languages
  3. Test with real documents using the free credits
  4. Integrate into your system and start scaling

DocsOCR was built for Brazilian companies that need OCR that simply works. No complex configuration, no lengthy integration projects — just a simple API that transforms documents into structured data.