Unify Logo Footer.svg
Unify Automations
Logo
PDF by UnifyApps

PDF by UnifyApps

Logo

3 mins READ

PDF by UnifyApps reads a PDF file page by page within an automation, yielding each page's extracted text as a data pill for downstream steps. Use it to process invoices, reports, and documents without manual extraction.

Overview

PDF by UnifyApps processes information by going through each page of a PDF and extracting the data. This extracted data can then be used throughout the automation using data pills.

Beyond reading text page by page, PDF by UnifyApps also includes actions for merging, extracting, and removing pages from a PDF, as well as:

  • Add Signature to PDF — adds a signature image to each page of a PDF file.

  • Split PDF File — iterates over a PDF, splitting it into multiple files by a specified page count.

Screenshot_2026-08-29_at_9.39.28_PM_1.png
Screenshot_2026-08-29_at_9.39.28_PM_1.png

Use Case

A practical scenario involves collecting invoices from Gmail and creating a Google Sheet with key information:

  1. PDF by UnifyApps extracts pages sequentially, leveraging Unify AI to split text and obtain essential information like Name, Source, and Amount.

  2. The data pills from PDF by UnifyApps can be used to add rows to Google Sheets files.

  3. This automation streamlines invoice collection and maintains clean records.

Use case example — extracting invoice data with PDF by UnifyApps and Unify AI

How to Parse a PDF

Step 1 — Add the PDF by UnifyApps node and select "Read text from the PDF file," then provide input for the PDF field.

Adding the PDF by UnifyApps node and selecting the action

Step 2 — Input can be a URL, a base64-encoded string, or a file object mapped from an upstream node.

Step 3 — Add an App and Action inside the node's iteration — for example, use Variable by UnifyApps to collect pages into a list.

Adding an App and Action inside the PDF node's iteration

Step 4 — Extract data using the output data pills, which contain:

  • Item — the text extracted from the current page

  • Page Index — the page number (zero-based)

  • Is First and Is Last — boolean indicators for the first and last page

Output data pills from the PDF by UnifyApps node

Note: A step must be added after the PDF by UnifyApps node to complete the automation. The automation cannot be deployed with an empty iteration.

Notes

Keep these behaviors in mind when using PDF by UnifyApps:

  • The node iterates page by page — you must add at least one step inside the iteration block before the automation can be deployed.

  • Input can be a URL, a base64-encoded string, or a file object from an upstream node — all three formats are accepted.

  • Use the Is First and Is Last flags to run setup or teardown logic (such as initializing a list or writing a final output) without adding an extra conditional branch.

  • Page Index is zero-based — page 1 in the document is Index 0 in the output.

  • This node extracts text only; it cannot parse images, tables, or embedded objects within PDF pages.

Pairing this node with an AI step (such as the Prompt node) lets you extract structured data from unstructured page text in a single automation pass.

FAQs

Can I provide a PDF file without a URL?

Yes — you can leverage file-type data pills within UnifyApps to pass the file directly into the reader.

Can this reader parse other file types like DOCX?

No — this node can only interpret PDF files.