Quick overview
This workflow receives invoice PDFs or images via a webhook, extracts text using OCR.space, parses key invoice fields and line items with deterministic rules, and then saves the results as a CSV file and optionally appends them to Google Sheets before returning the extracted data as JSON.
How it works
- Receives an HTTP POST webhook request containing an uploaded invoice file (PDF/image) in a binary field.
- Sends the uploaded file to the OCR.space API to convert the document into extracted text.
- Checks whether OCR.space processed the file successfully and returns a 422 JSON error response if OCR fails.
- Parses the OCR text to extract vendor, invoice number, invoice date, subtotal, tax, total, and line items, and generates flags for any missing fields or line-item calculation mismatches.
- Converts the extracted fields into a CSV file and writes it to disk.
- Optionally appends the extracted record to a Google Sheets worksheet.
- Returns the extracted invoice data (including flags and raw OCR text) as the webhook response.
Setup
- Create an OCR.space API key and add it to an n8n Header Auth credential (header name "apikey") used by the OCR request.
- Send invoices to the webhook endpoint by copying the production webhook URL from the "Invoice Upload" trigger and posting a file in the binary field named "invoice".
- Ensure n8n has write access to the configured output path (default: ./output/) or update the CSV file path to match your environment.
- (Optional) Add a Google Sheets OAuth credential, set the target Spreadsheet ID and sheet name, and ensure the header row matches the workflow’s output fields before enabling the Google Sheets append step.