Hi there,
Nexadata has turned PDFs into data for over a year. What it has not had is an experience worthy of the capability. Extraction worked, but everything around it asked too much of you, and the path from a document to a usable Dataset had more friction in it than it should have.
This release rebuilds that path end-to-end. PDF is now a Data Format handled like any other, the table review workflow has been redesigned around speed, and the rough edges are gone. The capability is not new. The experience is.
PDF is now a Data Format. It sits on the Connect Data step alongside Tabular and Spreadsheet, in the same radio group, so you create a PDF Dataset exactly the way you create any other one. Selecting PDF adds a Process PDF step to the wizard, which opens a dedicated workspace covering extraction, table review, and building.

Turning a document into a Dataset means making a decision about every table inside it, and a dense document detects more tables than you expect. Footers, address blocks, and signature panels all read as tables to an extraction engine. The new Table Selection stage is built so that getting through them is fast.

Every control does one thing, sits where you expect it, and gives you feedback immediately. That is the whole point of this release.
Every decision you make is recorded as a template. The next invoice, statement, or report in the same layout runs against it without repeating the review, because what the template depends on is the layout rather than the values. New numbers in the same tables are exactly the designed case.
That is why the review speed matters. A 600-page private equity quarterly report is a realistic input, and Nexadata stitches every table that spans a page break across the whole document. Curating it is real work the first time and close to free every quarter after.
The example in our documentation is a legal invoice, chosen because it is awkward in all the ways real documents are, but nothing here is specific to legal work. Supplier and freight invoices, bank and brokerage statements, utility bills, remittance advices, shipping manifests, lab results, insurance schedules, fund and board reports, regulatory filings. The source file showed you totals. The Dataset gives you the arithmetic behind them, per line, ready to aggregate any way you need.
๐ Setting Up a PDF Dataset โ
A column of hours or amounts read as text is effectively inert. You cannot sum it, average it, or group by it, which blocks the first thing most people want to do with extracted data.
The Dataset Builder now infers column data types automatically and lets you override the inference whenever it gets one wrong. A Hours column that came through as text becomes numeric, and the Group By with a Sum aggregation you were reaching for simply works.
The filter transformation could only express inclusion. Every operator was affirmative, so there was no clean way to say something as ordinary as "drop the rows where Account starts with 9." You had to build around the gap rather than through it.
Filters now express exclusion directly, with negated operators sitting alongside the existing ones. On extracted documents in particular, where subtotal and summary rows land in the middle of your detail, this is the difference between a clean Dataset and a fiddly one.
We are turning the "set it up once" idea into a first-class platform capability and broadening where your files can live:
Thanks for being part of the Nexadata journey!
The Nexadata Team