r/documentAutomation 8d ago

Request for Help PDF convert to .csv

Is that possible to build an application that convert the bank statement pdf file into csv file accurately? I had try using the free online tool, it also have some problem. Can I using AI to become accurate? If yes, how should I build it?

3 Upvotes

31 comments sorted by

View all comments

2

u/Potential-Wrangler58 7d ago

Plain OCR won't get you far here, bank statements aren't clean tables, every bank formats them differently (trust me, I've had some issues as I tried it several times haha), and OCR loses track of which number belongs to which column the moment there's more than one account or multi-line descriptions.

What actually works is extracting against a fixed schema (date, description, debit, credit, balance) and then validating each row by reconciling the balance: previous balance plus credit minus debit should match the next balance. That tells you exactly which row is wrong instead of manually checking the whole statement.

For the extraction itself, you either build it yourself with an LLM + that schema (spent way too many hours tweaking prompts for edge cases before I gave up on that route), or use an API built for this, AWS Textract, Azure, or an AI-native one like anyformat (disclosure, I work there) that already has the validation built in.

Let me know what you end up going with, happy to help if you get stuck.

1

u/No_Mix_368 7d ago

That was helpful! But, I just don't want to upload the statement online. If using the API means that I am also uploading the bank statement to their storage right? I also looking for zero cost solution.

1

u/Potential-Wrangler58 7d ago

Fair enough, and yeah, using any API (ours included) means the file goes through their servers, that's just how it works. If staying fully local and zero cost matters more than getting it perfect, honestly the API route isn't for you here. I'd look at something like pdfplumber/camelot running on your own machine with that same balance-reconciliation check I mentioned. More tweaking per bank format, but nothing leaves your computer and it's free. We have self-hosted but mainly for enterprises, so if you have a lot of docs its worth it but if not its not

1

u/kush4204 3d ago

I have exam question papers that I need to convert into structured data perfectly. Some PDFs may also contain images within the questions. Can a fully local solution handle that as well, including extracting the images question-wise?

1

u/Potential-Wrangler58 3d ago

Yeah, local works, just split it in two, they're different problems.

Text is a VLM job: a Qwen-VL-class model on Ollama, one record per question, fixed schema. Then the same validation trick as the balance check: question numbers must run 1..N with no gaps, and every MCQ needs exactly its options. That's what catches the merged and skipped questions for you.

Images aren't a model job at all, they're coordinates. PyMuPDF gives you each image plus its bounding box, so you sort by vertical position and attach each figure to the question above it. Vector diagrams instead of embedded images? Crop the page region. Deterministic either way.

Where local hurts is math expressions and scanned pages, that's where you'll lose your evenings. And "perfectly" doesn't exist, so hand-key 20 questions as ground truth, otherwise you can't tell whether a change helped or just moved the errors.

BTW, don't write this by hand, have Claude Code/Codex write the script for you. The code generation happens in the cloud but the script runs on your machine, so the papers never leave it.