Importing documents from email into Odoo
Turn vendor emails and printed PDFs into Odoo records with a template-driven pipeline built from Avunu's open-source import addons.
On this page
This document explains how Avunu's import addons fit together, so that a vendor quote, order confirmation or invoice that arrives by email, or that someone prints to PDF, becomes an Odoo record without retyping. Read it when you want to set up that pipeline, or when you need to add one more document type to it.
All of the modules live in the public avunu-odoo-addons repository, on the 18.0 branch. They are licensed AGPL-3.
How the pieces fit
Everything is built on the base_import_pdf_by_template module, which turns a document into text and pulls fields out of that text with per-line regular expressions. A template (base.import.pdf.template) targets one Odoo model, and each template line says which field to fill and which regex finds the value. Avunu's modules add sources, tooling and routing around that core.
| Module | What it adds |
|---|---|
base_import_pdf_by_template_engine | Two small extension points that the other modules plug into, plus fixes for two table-extraction bugs. |
import_from_email | plaintext and html extraction modes, a mail alias that creates EDI exchange records, and a generic processor that runs a template. |
import_via_xberg | An xberg extraction mode that parses a document into structured JSON, with a JSONPath step that narrows what a regex searches. |
import_preview | A saved sample document per template and a live preview of what each line matches. |
import_create_missing | Creates a missing linked record (a vendor, a product) during import instead of leaving the field unresolved. |
import_create_missing_xberg | Per-row JSONPath for the previous module. It installs itself when both of its dependencies are present. |
edi_import_router | One inbox for emailed and printed documents. It asks TypeSafe which document type each one is and files it accordingly. |
The flow for one document looks like this:
email to router alias ----+
+--> edi.import.router.document --> TypeSafe picks an exchange type
printed PDF (ERP Printer) + |
v
direct mail alias -----------------------------------------> edi.exchange.record
|
exchange type's import template
|
extract text, match lines, create the Odoo recordThe router is optional. If you only receive one kind of document from one sender, a direct alias on an exchange type is enough. The router earns its keep when many senders share one mailbox, or when some documents arrive as prints.
Install the modules
edi_import_router depends on import_from_email, which depends on base_import_pdf_by_template, the engine module and edi_core_oca. The router also needs mail and queue_job. Those dependencies come from OCA repositories: base_import_pdf_by_template from OCA/edi, edi_core_oca from OCA/edi-framework, and queue_job from OCA/queue. Install only what you need:
Email or printed documents routed automatically:
edi_import_router.Email to a fixed exchange type, no routing:
import_from_email.Structured PDF tables:
import_via_xberg.Template authoring tools:
import_preview.Creating vendors or products on the fly:
import_create_missing.
Install by name from the Apps menu, or from the command line:
python odoo/odoo-bin -c odoo.conf -d <DB_NAME> -i edi_import_router,import_preview --stop-after-initSome modules need Python libraries available to Odoo's interpreter. import_via_xberg declares jsonpath_ng and xberg, and edi_import_router declares typesafe_sdk and pypdf. Install them before you install the modules.
Note
The router classifies each document in a queue_job job, so the job runner has to be running. The OCA queue_job README says to load queue_job as a server-wide module (server_wide_modules = web,queue_job) and to restart Odoo after installing it on a database. Without a runner, routed documents never get classified.
Set up a template
A template is the unit of work. Build one per sender and document layout. Create the template with the target model, an extraction mode and, optionally, an auto-detect pattern.
Choose the extraction mode by where the document comes from:
| Mode | Source | Notes |
|---|---|---|
pypdf | A text-based PDF | Ships with base_import_pdf_by_template. |
plaintext | An email body or text file | Read back as UTF-8, unchanged. |
html | An HTML email body | Markup is stripped with Odoo's html2plaintext. Avunu adds a line break before every <div>, so a body built from nested divs keeps its line structure. |
xberg | A PDF (or other file) where tables matter | Produces a JSON document, described below. |
Each template line is either a header line, which yields one value for the document, or a lines line, which yields one value per row of a repeating table. Every lines line with a pattern becomes one column. Columns are zipped together by position into rows.
Warning
Columns are paired by position, not by content. A column that matches fewer times than its siblings is padded with blanks rather than dropped, so later columns stay aligned, but a pattern that skips a row will still put values on the wrong row. Check the preview (below) against a real sample before you trust a table.
Use xberg for tables
Flat text is a poor fit for a PDF table. The xberg mode parses the file into a structured document with real tables[].cells[][] grids and hands the template that JSON as its "page text". Install import_via_xberg and set the template's extraction mode to xberg.
Narrow with JSONPath, then refine with regex
pattern is always a plain regular expression, on every template. On an xberg template each line also has an optional Xberg JSONPath field. When you set it, the regex runs only against what the JSONPath selects, one match per line of text. When you leave it blank, the regex searches the whole JSON document.
Work in two steps: write the JSONPath until it selects the right values, then write a regex that refines each one.
The following examples use a table shaped like the one in the module's tests, with header cells QTY SHP, QTY B/O and ITEM:
| Xberg JSONPath | Selects | Pattern | Result |
|---|---|---|---|
| (blank) | The whole document | "lat":\s*"([\d.]+)" | The first match in the raw JSON, as on any other template. |
$.tables[0].cells[1][2] | The cell 201P/90-0013 | (\d+)$ | 0013 |
$.tables[0].cellsByHeader[*].ITEM | The ITEM value of every row | (\d+)$ | One refined value per row. |
To take a selected value as it is, use (.+) as the pattern.
Read a table by header
A raw cells grid forces you to hardcode column positions, and positions shift when a sender reorders columns. So every table also gets two derived views next to cells, which is never modified:
cellsByHeaderhas one object per data row, keyed by that column's header text. Row 0 is treated as the header and is not included.cellsByIndexhas the same rows keyed by the column index as a string ("0","1"). Use it for a column with a blank header, or when two columns share a header.
Not every table has a header row. A label and value table such as Order #: beside its number has none, and both derived views come back empty for it. Use cells for those.
If a header wraps across two lines in the PDF, the key contains a newline. Type it as a backslash followed by n inside a quoted key:
$.tables[1].cellsByHeader[*]['Part\n Number']\n, \t and \r are supported.
Warning
Against the raw cells grid, a slice followed by a wildcard (cells[1:][*][2]) descends into the strings and yields single characters. Write cells[1:].[2] instead, or use cellsByHeader and avoid the slice.
Tune the extraction
Two settings control how the file is parsed:
Xberg configuration on the template accepts a JSON object that is passed to the xberg library, for example
{"ocr": {"enabled": true}}. Leave it blank for the defaults. See the xberg documentation for the full schema.The system parameter
import_via_xberg.timeoutsets the extraction timeout in seconds. The default is300. A template's own configuration wins if it setstimeout_secs.
When a template could match any of several modes, Odoo tries every mode during auto-detection. The xberg mode is skipped when none of the candidate templates uses it, so you do not pay for a full parse on every upload.
Preview while you write
Install import_preview. It adds a Sample Data tab to the template. Paste a sample document there, or upload a file and let the template's own extraction fill the field in. Uploading is the right choice for a PDF or an xberg template, because it stores exactly the text your patterns will see.
With a sample in place:
Each line dialog shows a Preview under
Patternthat updates as you type. It runs the same extraction and conversion code as a real import, so dates, numbers and lookups behave the same way. Alinespreview shows up to 25 values, and an invalid regex shows a warning instead of an error.On an
xbergtemplate, the Xberg JSONPath field has its own preview, listing each numbered match. Check the narrowing step there before you write the regex.The Sample Data tab shows a Preview of the whole template: the assembled header values and table rows. Misaligned columns show up here first.
On a header line with a Search field, a lookup that finds nothing is called out. This is the most common template mistake, because the import otherwise falls back to the raw text silently.
One caveat: if you paste sample data and open an existing line without saving the template first, the first render can be stale. It corrects itself on your next keystroke in the dialog.
Route email into Odoo
The mail side has two options. Both file the document as an edi.exchange.record, which is a tracked, retryable EDI record.
Option 1: a direct alias
Use this when a mailbox always carries one document type.
Create the input
edi.exchange.type. Set Processor to "EDI Input Process: Import by Template" and Import Template to your template.Enable
quick_execon the type if you want processing to start immediately.import_from_emailmakes this checkbox visible on input types, where the stock view hides it.Create a
mail.aliaspointing atedi.exchange.record, and set its defaults to a Python dict literal naming the backend and the exchange type code:
{'backend_id': <BACKEND_ID>, 'type_code': '<TYPE_CODE>'}backend_id and type_code are keys in the alias's defaults text, not fields you will find on the record. A message routed there without both keys raises an error instead of creating a half-valid record.
The raw HTML body is stored unchanged as the exchange file and the record starts in the input_received state, since the whole document has already arrived. Interpretation is the template's job, so use an html or plaintext template. Every inbound email becomes its own record, even when a mail client threads several forwards under one subject.
The generic processor reads the template from the exchange type, extracts the text, and creates the target record through the same form-driven engine a manual upload uses. It then links the record to the exchange record.
Note
The generic processor creates records and nothing else. A document that needs fuzzy partner or product matching, safe updates to an existing record, or de-duplication deserves its own dedicated processor on its exchange type. Both kinds can coexist on one backend.
Option 2: the import router
Use edi_import_router when several senders share one inbox, when staff forward vendor mail by hand, or when a document only exists as a printout.
In Settings, enter the TypeSafe API key, or set the
TYPESAFE_API_KEYenvironment variable. TypeSafe is a hosted AI service that the router calls through thetypesafe_sdkPython package, so Odoo needs outbound internet access. With no key, every document waits for manual review.Create an input exchange type and template for each document type, as in option 1.
Open Import Routing, then Import Routers (under the EDI menu), and create a router. Give it an email alias, a responsible user, and one target per exchange type.
Write each target's Description to say what tells that document apart: the sender, the subject or title, distinctive text. TypeSafe chooses between these descriptions, so name the difference explicitly when two vendors send from the same address.
A target can be limited to email or to printed documents. Whatever you choose, a document is only offered to a type whose template can read it: an email body needs an html or plaintext template, and a PDF needs a PDF template. Pointing a printer-only target at a body template is rejected.
Each arriving document becomes an edi.import.router.document and a background job (queue_job) classifies it, so neither the mail gateway nor the printer waits on the API. The job sends TypeSafe up to 8000 characters of the email body or PDF text, along with the subject, the sender and the original sender found in a forwarded header. The router's Question field holds the instruction text, and its default already explains that forwarded copies of a vendor document are that document.
The outcome depends on two thresholds on the router:
| Setting | Default | Meaning |
|---|---|---|
| Confidence threshold | 0.8 | A document is filed automatically only when TypeSafe is at least this sure. |
| Ignore threshold | 0.98 | A document is ignored only when TypeSafe answers "none" with at least this much confidence. |
The ignore threshold is stricter on purpose, because a wrongly ignored order costs more than an extra review item. A confident match is filed with the same create_record() call that import_from_email uses. A document that TypeSafe doubts, or that fails, lands in review with a to-do for the responsible user. Rate limits and connection errors are retried up to three times, 60 seconds apart, before a person is asked.
Reviewers open Routed Documents, which is filtered to documents needing review by default. Each document shows its state (pending, routed, review, ignored or error), TypeSafe's confidence and per-target probabilities, and the resulting exchange record. Three buttons handle the rest: Dispatch files it under the target you picked, Re-classify asks again, and Ignore closes it.
Two access groups exist. Import router reviewer can read the router and work on documents. Import router: printer upload can attach a PDF to a router and nothing else.
Forwarding from Gmail
A forwarding rule matches the router alias only if Odoo sees the alias as a Delivered-To recipient. When mail arrives through the Cloudflare module, the envelope recipient has to be added even if the message already carries a Delivered-To from the forwarder. The Avunu mail_cloudflare module in the same repository handles that inbound path.
Printed documents
A quote that only exists on a web page can be printed to PDF and uploaded by ERP Printer.
Create an Odoo user in the Import router: printer upload group, with an API key.
On the ERP Printer profile, set the backend to
OdooJsonRpc(Odoo 18), the Odoo model toir.attachment, Attach to model toedi.import.router, and Attach to record ID to the router's ID.
The router form shows these values in a read-only line. The upload becomes a print document and is classified like an email. Files with identical contents are ignored, so a retry from the printer's outbox never files a document twice.
Create missing records
By default, if a line's search finds no vendor or product, the import falls back to the line's fixed value or the raw text. import_create_missing can create the missing record instead.
On a template line whose Field is a many2one with a Search field set, check Create New Document if Not Found.
Open the New Document Values tab. It fills itself with every field the new record requires. Fields with no Odoo default appear as empty
Variablerows, and fields with a default appear asFixedrows already set to that default.Fill the empty rows, and adjust any default you do not want. Add rows for anything else the record should carry.
Import a document whose value will not be found.
Each row sets one field and has a Type:
Fixed: always the same value, entered with an input that matches the field's type.
Variable: extracted with a regex, evaluated once per table row (or once for a header line), in the same order as the line's own values.
Odoo Default: left out of the create, so Odoo applies its own default. This suits a field that means "now".
The search value that missed is written into the new record's search field, so importing the same source again finds the record instead of duplicating it.
Rows are paired with the line's rows by position. If a row's pattern produces a different number of matches than the line itself, no record is created for any row of that line, and the reason is logged. Positions are never guessed.
If the field you are filling is a one2many, such as a vendor's bank accounts or a product's vendor list, its Type, Value and Pattern are replaced by a Row Values table for the sub-record's own fields. A product line always offers a Vendors row, because a product created without one can never be found by vendor code again.
A row whose required field is still empty is highlighted and labelled Missing, and the tab shows a banner. The warning is advisory. The line still saves, so you can build a template over several sittings. Add Required Fields appends what is missing without touching rows you edited.
Note
Records are only created during a real import, from a manual upload or from the EDI intake. A preview, whether the line dialog or the standalone preview wizard, never creates anything.
On an xberg template, import_create_missing_xberg adds a JSONPath field to each Variable row. Use the same table the line reads, for example $.tables[1].cellsByHeader[*]['Description'] to fill a name from the Description column of the row where the line found a part number. A row's pattern is optional here, and when set it runs against each JSONPath match on its own. A blank match stays blank at its position, so one empty cell never shifts the rows after it. A Variable row that has a pattern but no JSONPath of its own searches the whole JSON document, and the warning banner names it.
Troubleshooting
| Symptom | Likely cause |
|---|---|
| A message to a direct alias raises an error. | The alias defaults lack backend_id or type_code. |
| An EDI record never processes. | The type has quick_exec on but the record is not in input_received. Records created by these modules already are. |
| HTML tags show up in the text your patterns see. | An HTML email body is being read by a plaintext template. Use html. |
| Every document goes to review. | No TypeSafe API key is set, or confidence is below the router's threshold. Read the document's error message. |
| A document says no target can read it. | The targets' templates cannot read that source, such as only body templates for a printed PDF. |
| A creation is skipped for a whole table. | A New Document Values column matched a different number of times than the line. Fix the pattern. |
| A record was not created and no error is shown. | Turn on developer mode, open Settings, Technical, Logging and filter by Path import_create_missing. Entries are written on their own database cursor so a rolled-back import does not erase them. |
| An xberg preview says to add a sample. | The template has no sample document yet. Upload one on the Sample Data tab. |
What this document does not cover
This document does not describe the base module's own screens, such as the manual upload wizard and the template list, because that module is not part of Avunu's repository. Menu and field labels here come from the Avunu modules' own views. Test commands for each module are in its README, and run with --test-enable --test-tags /<MODULE_NAME>; Testing Custom Odoo Modules explains the options.
Related documents
Sources
- github.com/Avunu/avunu-odoo-addons
- github.com/Avunu/avunu-odoo-addons/tree/18.0/edi_import_router
- github.com/Avunu/avunu-odoo-addons/tree/18.0/import_from_email
- github.com/Avunu/avunu-odoo-addons/tree/18.0/import_via_xberg
- github.com/Avunu/avunu-odoo-addons/tree/18.0/import_preview
- github.com/Avunu/avunu-odoo-addons/tree/18.0/import_create_missing
- github.com/Avunu/avunu-odoo-addons/tree/18.0/base_import_pdf_by_template_engine
This article is in the public domain (CC0 1.0), code samples included. Use it however helps you.