The production arrives as one PDF of four thousand scanned pages with no text layer, a filename of PROD0001.pdf, and a deposition in nine days.
The short answer
A scanned production is unsearchable until something reads it, which makes OCR the first step rather than an enhancement. A workflow uses the Read Details from File step to read the pages, the Split Document step to break a combined scan into the documents it actually contains, and the Rename step to name each piece by its date, type and parties. The Move step then files them under the matter. What this buys is the ability to search a production by what the documents say — the difference between having the documents and being able to work them.
Steps this uses
Before and after
As they arrive
After the workflow
One combined scan split into the documents it contained, each named by what it is and when it was written.
Setting it up
This is the sentence. Send it to the builder and the steps below appear on a canvas, wired and named, for you to change before anything runs.
When a scanned production is uploaded, read the pages, split the scan into the documents it actually contains, rename each one by its date, type and parties, and file them under the matter in Discovery.
Nothing downstream works on an image. The Read Details from File step reads the pages first, and every later step — splitting, naming, searching — depends on that having happened.
A production is rarely one document. The Split Document step separates it at the boundaries between documents, so a four-thousand-page PDF becomes the letters, invoices and memoranda it was always made of.
The Rename step writes what the document is, when it was written and who it was between. That is what makes a chronology assemblable by sorting rather than by reading.
The production as received is the record of what was produced. Work on copies, and keep the original file exactly as it arrived, with its filename and its pagination intact.
Every firm already has the discovery. It is on the drive, complete, backed up, and effectively unavailable, because finding the one letter that matters means opening a four-thousand-page PDF and scrolling. The work of litigation is retrieval under time pressure, and a production without a text layer converts that into reading. OCR is not a convenience here; it is the thing that makes the set usable at all.
Most of what a case turns on is sequence: who knew what, and when. A production named by date, document type and parties sorts into a chronology by filename alone, which means the timeline is a by-product of filing rather than a separate week of work. Names that start with a Bates number sort into the order the producing party scanned things, which is an order chosen by your opponent.
It is not an eDiscovery platform. There is no Bates numbering, no load file, no review coding, no privilege log, no production export. A large matter with a review team and a defensible process needs a dedicated platform, and this does not replace one. Where it earns its place is the smaller matter, and the productions a firm receives rather than produces.
FAQ
Run OCR over it, which reads the text off the scanned images and makes the pages searchable by what they say. Until that happens the documents are pictures, and no filing structure helps — you can find the folder but not the letter. OCR has to come before splitting, naming or anything else.
Yes. A production that arrives as one long PDF can be split at the boundaries between documents, so the letters, invoices and memoranda inside it become separate files that can be named and sorted. This is normally the difference between a production you can work and one you can only scroll.
Keep the Bates number, but do not lead with it. A filename starting with a Bates number sorts into the order the producing party scanned things. Leading with the date and naming the document type and parties gives you a chronology by sorting, which is what the case usually turns on.
No, and it is worth being clear about it. There is no Bates stamping, load file, review coding or privilege logging. A large production with a review team and a defensible process belongs in a dedicated platform. This is for making received productions searchable and for matters that do not justify one.
Read, split and named by what each document actually is.
5 GB free · No credit card required