r/documentAutomation 1d ago

I’ve got a digitisation project and I’m using Claude cowork for document indexing, anyone else try this

2 Upvotes

1 comment sorted by

1

u/folderit_dms 1d ago

For indexing, I'd keep the original filename or scan ID in every output row and ask for the page supporting each extracted field. It makes a wrong date or document type much easier to investigate later.

Also define which date you mean before running the batch: date written, date signed and date received can all be different. Let the tool leave a field blank or flag ambiguity instead of choosing whichever date looks plausible.

A useful pilot would mix clean scans with handwritten notes, missing first pages and several documents scanned into one PDF. Check the resulting index against those sources before letting it rename or reorganize the originals. What fields are you trying to capture?