Information Extraction with AI (Azure Content Understanding)

Extract information from documents

Azure Content Understanding

LEARN what Azure Content Understanding is and how it works

Azure Content Understanding goes beyond simple OCR text extraction. It follows a schema-based extraction of fields and values.

It follows this workflow

  1. Ingest content
  2. AI-powered analysis: The services analyzes the content through combination of
    • OCR
    • speech recognition
    • NLP/NLU
    • multimodal AI models
  3. Structured output - returns for example JSON

Understand analyzers

Un analyzer es una unidad de Azure Content Understanding la cual:

  • toma input
  • aplica analisis mediante IA
  • produce un resultado estructurado

Una vez estructurado un analyzer se encarga de que un schema sea reusado de manera consistente para cada request. Un analyzer produce resultados predecibles en JSON.

Azure Content Understanding ofrece prebuilt analyzers o se pueden customizar. A alto nivel:

  1. Eliges o creas un analyzer
  2. El analyzer incluye un schema que define los campos y su estructura
  3. Adjuntas contenido para ser analizado
  4. El servicio aplica el schema
  5. Recibes JSON de vuelta que concuerda con el schema

Example

With Azure Content Understanding if we declare the following schema

  • Vendor name
  • Invoice number
  • Invoice date
  • Customer name
  • Custom address
  • Items - the items ordered, each of which includes:
    • Item description
    • Unit price
    • Quantity ordered
    • Line item total
  • Invoice subtotal
  • Tax
  • Shipping Charge
  • Invoice total

And apply this to a document

document example

It will extract the fields with their nested values

  • Vendor name: Adventure Works Cycles
  • Invoice number: 1234
  • Invoice date: 03/07/2025
  • Customer name: John Smith
  • Custom address: 123 River Street, Marshtown, England, GL1 234
  • Items:
    • Item 1:
      • Item description: 38” Racing Bike (Red)
      • Unit price: 1299.00
      • Quantity ordered: 1
      • Line item total: 1299.00
    • Item 2:
      • Item description: Cycling helmet (Black)
      • Unit price: 25.99
      • Quantity ordered: 1
      • Line item total: 25.99
    • Item 3:
      • Item description: Cycling shirt (L)
      • Unit price: 42.50
      • Quantity ordered: 2
      • Line item total: 85.00
  • Invoice subtotal: 1409.99
  • Tax: 140.99
  • Shipping Charge: 35.00
  • Invoice total: 1585.98

Azure Content Understanding In Foundry

You candeploy ACU in Foundry in

Discover > Services > Azure Content Understanding

azure content understanding

test ACU example

result json example

Azure Content Understanding SDK

Install

python -m pip install azure-ai-contentunderstanding

import os
from azure.ai.contentunderstanding import ContentUnderstandingClient
from azure.core.credentials import AzureKeyCredential

endpoint = os.environ["FOUNDRY_ENDPOINT"]
key = os.environ["FOUNDRY_KEY"]

client = ContentUnderstandingClient(endpoint=endpoint, credential=AzureKeyCredential(key))

# 1) start analysis with analyzer id + inputs
analyzer_id = "prebuilt-invoice"
inputs = [
    {"url": "https://github.com/Azure-Samples/azure-ai-content-understanding-python/raw/refs/heads/main/data/invoice.pdf"}
]

# 2) wait for the Long Running Operation (LRO) to complete
poller = client.begin_analyze(analyzer_id=analyzer_id, inputs=inputs)  # starts LRO
result = poller.result()  # waits for completion (polling handled by SDK)

# 3) read structured fields + markdown
# The result typically includes extracted "fields" and "markdown" per input content item.
for content in result.contents:
    print(content.markdown)
    print(content.fields)

resulting output

{
	"status": "Succeeded",
	"result": {
		"analyzerId": "prebuilt-invoice",
		"apiVersion": "2025-05-01-preview",
		"contents": [
			{
				"markdown": "# INVOICE\n\nCONTOSO LTD.\n\nContoso Headquarters\n123 456th St\nNew York, NY, 10001\n\nINVOICE: INV-100\n\nINVOICE DATE: 11/15/2019\n\nDUE DATE: 12/15/2019\n\nCUSTOMER NAME: MICROSOFT CORPORATION\n",
				"fields": {
					"CustomerName": {
						"type": "string",
						"valueString": "MICROSOFT CORPORATION",
						"confidence": 0.95,
					},
					"InvoiceDate": {
						"type": "date",
						"valueDate": "2019-11-15",
						"confidence": 0.994,
					}
                }
            }
        ]
    }
}

Extract information from audio and video

Azure Content Understanding supports both audio and video files.

Extract structured data from audio

Azure Content Understanding provides transcriptions, summaries and other insights from audio files.

You define a schema

  • Caller
  • Message summary
  • Requested actions
  • Callback number
  • Alternative contact details

Process a voice mail

(imagine this is an audio file)

Hi, this is Ava from Contoso.

Just calling to follow up on our meeting last week.

I wanted to let you know that I've run the numbers and I think we can meet your price expectations.

Please call me back on 555-12345 or send me an e-mail at Ava@contoso.com and we'll discuss next steps.

Thanks, bye!

Then Azure Content Understanding processes and gives us a response

  • Caller: Ava from Contoso
  • Message summary: Ava from Contoso called to follow up on a meeting and mentioned that they can meet the price expectations. They requested a callback or an email to discuss next steps.
  • Requested actions: Call back or send an email to discuss next steps.
  • Callback number: 555-12345
  • Alternative contact details: Ava@contoso.com

Extract structured data from video

Azure Content Understanding also supports video analysis the same way. You give it a schema, process a video and it results the JSON

Process audio/video through SDK

import os
from azure.ai.contentunderstanding import ContentUnderstandingClient
from azure.core.credentials import AzureKeyCredential

# Endpoint and key for your Foundry resource
endpoint = os.environ["FOUNDRY_ENDPOINT"]  # e.g., "https://<resource>.services.ai.azure.com/"
key = os.environ["FOUNDRY_KEY"]

client = ContentUnderstandingClient(
    endpoint=endpoint,
    credential=AzureKeyCredential(key)
)

# Choose a prebuilt analyzer for audio
# (The documents module lists examples like prebuilt-audioSearch / prebuilt-videoSearch.)
analyzer_id = "prebuilt-audioSearch"

# Provide an input audio file (URL shown here; you can swap in your own accessible media URL)
inputs = [
    {"url": "https://<your-host>/samples/voicemail.wav"}
]

# Start analysis (asynchronous long-running operation)
poller = client.begin_analyze(analyzer_id=analyzer_id, inputs=inputs)

# Wait for completion (SDK polls under the hood)
result = poller.result()

# Inspect the structured output (JSON-like objects)
for content in result.contents:
    # Some analyzers may return a transcript and/or extracted fields depending on the analyzer and schema
    print("=== MARKDOWN / TRANSCRIPT (if provided) ===")
    print(getattr(content, "markdown", None))

    print("\n=== EXTRACTED FIELDS ===")
    print(getattr(content, "fields", None))