Python

Python examples

Six examples ready to copy, from the first call to error handling, with the hosted sample certificate (fictitious data).

Before you start

Three steps, once.

  1. Create an API key

    In the panel, under API Keys.

  2. Put the key in the environment

    In the DOCSOCR_API_KEY variable, which the examples read. The prices example needs no key.

  3. Save and run

    Save each example under the name shown above its code. Its first lines say what it needs to run.

Your first call

Sends the hosted sample certificate (fictitious data) and prints the JSON answer.

first_call.py
"""Your first call: extract the hosted sample certificate (fictitious data).

Needs the requests package and DOCSOCR_API_KEY in the environment.
"""
import os
import uuid

import requests

response = requests.post(
    "https://api.docsocr.com/api/v1/documents/birth-certificate",
    headers={"Authorization": f"Bearer {os.environ['DOCSOCR_API_KEY']}"},
    json={
        "imageType": "url",
        "imageUrl": "https://docsocr.com/samples/certidao-nascimento-exemplo.jpg",
        "requestId": str(uuid.uuid4()),  # one id per document
    },
    timeout=120,  # above the API's 90 s extraction budget
)
answer = response.json()
if response.status_code != 201 or not answer["success"]:
    raise SystemExit(f"{response.status_code} {answer.get('errorCode', '')} {answer.get('message') or answer.get('error')}")
print(answer["data"]["dados_pessoais"]["nome_completo"])

A local file in base64

Reads an image from your machine and sends it in base64, with no public URL needed.

local_file.py
"""Extract a certificate from a local file, sent in base64.

Needs the requests package and DOCSOCR_API_KEY in the environment.
"""
import base64
import os
import uuid

import requests

with open("certificate.jpg", "rb") as file:
    image_base64 = base64.b64encode(file.read()).decode("ascii")

response = requests.post(
    "https://api.docsocr.com/api/v1/documents/birth-certificate",
    headers={"Authorization": f"Bearer {os.environ['DOCSOCR_API_KEY']}"},
    json={
        "imageType": "base64",
        "imageBase64": image_base64,
        "requestId": str(uuid.uuid4()),  # one id per document
    },
    timeout=120,  # above the API's 90 s extraction budget
)
answer = response.json()
if response.status_code != 201 or not answer["success"]:
    raise SystemExit(f"{response.status_code} {answer.get('errorCode', '')} {answer.get('message') or answer.get('error')}")
print(answer["data"]["dados_pessoais"]["nome_completo"])

An image outside our standard

With resizeImage: true, an image outside our size standard is resized on our side for 1 credit more, instead of refused.

resize.py
"""Extract an image outside our standard: resizeImage brings it to the
standard first, for one extra credit (the 800x640 sample is too small).

Needs the requests package and DOCSOCR_API_KEY in the environment.
"""
import os
import uuid

import requests

response = requests.post(
    "https://api.docsocr.com/api/v1/documents/birth-certificate",
    headers={"Authorization": f"Bearer {os.environ['DOCSOCR_API_KEY']}"},
    json={
        "imageType": "url",
        "imageUrl": "https://docsocr.com/samples/certidao-nascimento-exemplo-800x640.jpg",
        "requestId": str(uuid.uuid4()),  # one id per document
        "resizeImage": True,
    },
    timeout=120,  # above the API's 90 s extraction budget
)
answer = response.json()
if response.status_code != 201 or not answer["success"]:
    raise SystemExit(f"{response.status_code} {answer.get('errorCode', '')} {answer.get('message') or answer.get('error')}")
print(answer["data"]["dados_pessoais"]["nome_completo"])
print("resized:", answer.get("imageResized", False), "credits:", answer["creditsCharged"])

The fast engine

Asks for the fast engine: it is tried first, at its own price, and the standard engines answer when it cannot. The answer names the engine that answered and what it cost.

fast.py
"""Ask for the fast engine: it is tried first, at its own price, and the
standard engines follow when it cannot answer. The answer names the engine
that answered and what it cost.

Needs the requests package and DOCSOCR_API_KEY in the environment.
"""
import os
import uuid

import requests

response = requests.post(
    "https://api.docsocr.com/api/v1/documents/birth-certificate",
    headers={"Authorization": f"Bearer {os.environ['DOCSOCR_API_KEY']}"},
    json={
        "imageType": "url",
        "imageUrl": "https://docsocr.com/samples/certidao-nascimento-exemplo.jpg",
        "requestId": str(uuid.uuid4()),  # one id per document
        "engine": "fast",
    },
    timeout=120,  # above the API's 90 s extraction budget
)
answer = response.json()
if response.status_code != 201 or not answer["success"]:
    raise SystemExit(f"{response.status_code} {answer.get('errorCode', '')} {answer.get('message') or answer.get('error')}")
print("engine:", answer["engine"], "credits:", answer["creditsCharged"])

Errors and retries

Handles every answer and retries with the same requestId: for 15 minutes, a repeat of a finished request gets the same answer at no charge, and a repeat of one still running gets a 409, to wait.

errors.py
"""Extract a certificate, handling every answer, with retries that never charge twice.

Each retry sends the same requestId: a repeat of a finished request gets its
kept answer at no charge, and a repeat of one still running is told to wait (409).
Needs the requests package and DOCSOCR_API_KEY in the environment.
"""
import os
import time
import uuid

import requests

ENDPOINT = "https://api.docsocr.com/api/v1/documents/birth-certificate"
# No engine could answer: nothing was charged, and a retry may succeed
RETRY_LATER = {"EXTRACTION_BUSY", "EXTRACTION_TIMEOUT", "EXTRACTION_UNAVAILABLE"}
# Still in progress, too many requests, the service restarting
RETRY_STATUS = {409, 429, 502, 503, 504}


def extract(image_url, attempts=5):
    request_id = str(uuid.uuid4())  # one id per document, the same on every retry
    for attempt in range(attempts):
        backoff = 2**attempt  # seconds: 1, 2, 4, 8
        try:
            response = requests.post(
                ENDPOINT,
                headers={"Authorization": f"Bearer {os.environ['DOCSOCR_API_KEY']}"},
                json={"imageType": "url", "imageUrl": image_url, "requestId": request_id},
                timeout=120,  # above the API's 90 s extraction budget
            )
        except requests.RequestException:  # a timeout or a dropped connection
            time.sleep(backoff)
            continue
        is_json = response.headers.get("Content-Type", "").startswith("application/json")
        answer = response.json() if is_json else {}
        if response.status_code == 201 and answer.get("success"):
            return answer
        # Wait as long as the answer asks (retryAfter); a limit that resets
        # later, such as a daily quota, is not worth waiting for
        wait = answer.get("retryAfter") or backoff
        retry = answer.get("errorCode") in RETRY_LATER or response.status_code in RETRY_STATUS
        if retry and wait <= 60:
            time.sleep(wait)
            continue
        # 400 the body, 401 the key, 402 NOT_ENOUGH_CREDITS, 422 the image
        # (its errorCode says what to fix), DOCUMENT_NOT_RECOGNIZED: fix it, don't retry
        raise SystemExit(f"{response.status_code} {answer.get('errorCode', '')} {answer.get('message') or answer.get('error')}")
    raise SystemExit("No answer after retries: try again later")


answer = extract("https://docsocr.com/samples/certidao-nascimento-exemplo.jpg")
print(answer["data"]["dados_pessoais"]["nome_completo"], "credits:", answer["creditsCharged"])

Prices

Reads the price of an extraction in credits, per engine and for resizeImage. No key needed.

prices.py
"""The price of an extraction, in credits, per engine and for resizeImage.

A public endpoint: no key needed. Needs the requests package.
"""
import requests

response = requests.get("https://api.docsocr.com/api/v1/documents/prices", timeout=30)
response.raise_for_status()
for name, credits in response.json().items():
    print(f"{name}: {credits} credit(s)")

Ready to start?

Create your account, get free credits and run the first example with your key.