Automate the obvious. Keep humans in control.

DocuFlow AI pulls data out of invoices, contracts and purchase orders. I designed the review experience: how people check what the AI extracted, fix what's wrong, and send it on for approval.

Product

DocuFlow AI (Neural Connect)

My role

Product designer, extraction and review

Team

2 designers, PM, ML lead, engineers

Status

Designed and prototyped, not launched

The DocuFlow review screen: question groups and AI assistant on the left, the invoice in the middle, and the review panel on the right.
The review screen. The document sits in the middle. Fields that need a person are listed on the right, and values that look right stay out of the way.

The problem

Finance teams typed every field by hand, then a second person checked it all again. Checking took almost as long as typing.

What I designed

A review flow where the AI fills in the fields and people only spend time on values that are uncertain or important. The source is always one click away.

Where it stands

Designed and prototyped with the ML lead and engineers. It hasn't launched, so I've defined how we'd measure it, starting with the AI mistakes people catch.

01The problem

Every field was typed by hand,
then checked again by someone else.

Invoices, contracts and purchase orders came in by email or through shared folders. Someone opened each one and typed the data into Excel or another system. A second person checked it against the document, and only then did it go for approval.

ReceiveBy email or a shared folder
Type it inRead the document, enter every field
Check itA second person reads it all again
Send onEmail it to the approver
ApproveManager reviews and signs off
SaveFiled in a folder or system

Six manual steps. The two red ones are the same work done twice, and that's where most of the time went.

What this cost the team

  • Time.One invoice took about 10 minutes to type and 5 to check. A reviewer got through roughly 30 documents a day.
  • Repeated work.Every document was read twice, once to enter it and once to check it.
  • No shared status.Updates lived in email threads and folders, so managers couldn't see what was pending.
  • Late errors.Mistakes usually surfaced at approval, after two people had already handled the document.
Automating data entry was the easy part. The hard part was helping people know when to trust the result.

02My role

I owned extraction and review,
from upload to approval.

I ran the research, made the flows and prototypes, and worked every day with the PM, the ML lead and the engineers. Another designer owned settings and the audit trail screens. We shared one set of patterns so the product felt consistent.

  • Extraction and review.How a document moves from upload to approval.
  • Confidence.How people read AI confidence and act on it.
  • Checking and correcting.How people verify a value and fix it.
  • Handoffs.How work moves between reviewers, leads and approvers.
  • Dashboards.What each role needs to see.
  • UI and design system.Components, patterns and prototypes.

03Research

I sat with people while they did the work.

Before designing anything, I wanted to see where the time went, where mistakes happened, and what made people trust or doubt the data.

Reviewers5 people, 2 sessions each

Entered data and checked it against the original document.

Team leads2 people

Assigned work and helped reviewers when they got stuck.

Operations manager1 person

Tracked errors and needed to know what changed, and who changed it.

What I found

Checking took almost as long as typing

People re-read the whole document to check the data, even when most of it was obviously right.

So the screen should point people to the few values that need a look.

They re-read everything because they couldn't see the source

The checker only had the spreadsheet. To verify one value they had to find it again in the document, so they re-read all of it to be safe.

So every value needs to link back to where it came from on the page.

Some fields matter more than others

Nobody worried about a fax number. Everyone double-checked totals, tax numbers and bank details, whatever the system said.

So importance and confidence are different things, and the design has to handle both.

Managers cared about changes, not values

The operations manager didn't need to see every value. They needed to know what was changed, by whom, and why.

So every correction is recorded, and the original value is kept.

04Key decisions

Let the AI handle the easy cases.
Let people make the decisions.

People wanted less checking, but not at the risk of missing a wrong total. So the real question was how the screen tells someone where to look.

01 — Show confidence where it changes what people do

Decision

Fields fall into three bands. High confidence (90% and up) fills in on its own. Medium (70–89%) and low (under 70%) wait for a person, with the low ones first.

Why

Research showed people re-reading whole documents when most of it was right. Sending them straight to the few fields that need a look, 5 out of 14 on the invoice above, is where the time saving comes from.

Extracted fields marked high, medium and low confidence, so attention goes to the uncertain ones.

02 — Keep the document next to the data

Decision

The document stays on screen the whole time. Selecting a field highlights it on the page, and the edit drawer opens beside the document, not over it.

Why

The second check existed because the checker only had a spreadsheet. With the source right there, there's no reason to re-read everything.

The document and the extracted values side by side, so a value can be checked against its source.

05Trade-offs

What I decided not to do.

  • Full automation. The AI is sometimes wrong, and in finance a wrong total is expensive. People stay in the loop for anything uncertain.
  • A score on every field. Fourteen percentages on one screen is noise. Scores appear where they help someone decide what to check.
  • A separate source view. Early versions had Edit and Source tabs in the drawer. The document was already on screen, so the Source tab only repeated it. I removed it and put a small preview of the source line right above the input.
  • A screen for each step. Extract, check and approve happen in one place, so people can finish a document without moving between screens.

06The workflow

How the workflow changed.

Before, every field was typed and then checked. Now the AI fills in the fields and the screen sends people to the few that need attention.

The workflow before and after: manual typing and repeated checking on one side, guided review on the other.

WorkflowUpload → Extract → Review → Approve → Complete

An uncertain field being corrected and marked as reviewed.
Fixing an uncertain valueThe reviewer compares the value with the source, corrects it if needed, and marks it reviewed. The AI's original value is kept for the record.

07Screen anatomy

The review screen, part by part.

Reviewers spend most of their day here, so every part had to earn its space.

The review screen with its seven regions numbered and labelled.
  1. 1Product railDashboard, inbox, review queue and templates. Labels help new users learn their way around.
  2. 2Question groupsWhat the AI needs to find. Also works as a filter, so reviewers can focus on one group at a time.
  3. 3AI assistantExplains why a field was flagged and can suggest a value. It never makes the final call.
  4. 4DocumentThe original file. Highlights show which values need attention.
  5. 5Confidence popoverThe score, the reasons behind it, and a way to review the field.
  6. 6Review panelProgress, the confidence breakdown, and the fields still waiting for a decision.
  7. 7Edit drawerOpens beside the document, with a preview of the source line above the input, so people compare and type in one place.

08Dashboards

Different roles ask different questions.

The same document passes through reviewers, team leads and managers. Reviewers want to know what to work on next. Team leads want to see where work is stuck. Managers want to know what changed and how many errors got through.

Below is the reviewer view. It answers one question: what do I need to work on? That means what's assigned to me, what needs review, what's urgent, and what has been waiting too long. The goal is that nobody has to ask a lead what to do next.

The reviewer dashboard showing assigned documents, what needs review and what is overdue.

09Discover questions

Let the documents suggest the questions.

Problem

Before the AI can answer anything, someone has to write the questions, like "What is the invoice number?". Doing that by hand was slow, and with a new document type people had to guess what it even contained.

Decision

I added Discover questions to Question groups. After uploading documents, the user clicks the wand icon, the AI reads the batch, and it suggests questions it can answer from those documents. The user picks the ones they want and adds them to a group.

Why

It cuts setup time and gets people to useful answers sooner. Nothing is added without the user choosing it, because one wrong question means wrong data on every document after it.

Q4 vendor invoices 12 documents Discover questionsSuggest questions from the documents in this folder PS
An icon on its own is easy to miss, so the wand has a tooltip and is highlighted the first time someone opens a folder that has documents but no questions.
1. Pick documentsAll in the folder, or a selection
2. AI reads themProgress per document, can run in the background
3. Choose questionsReview, edit the wording, add to a question group
Each suggestion shows how many documents it was found in and a real value from those documents, so people can judge it without opening a file. Common questions start ticked, rare ones don't, and questions already in the group are greyed out. If some files couldn't be read, the modal says so instead of quietly suggesting less.

10Audit trail

Every change has a name, a time
and the original value.

The operations manager's question from research was simple: what changed, and who changed it? In finance that isn't a nice-to-have. Auditors ask it too.

The other designer owned the audit trail screens. My part was deciding what the review flow records. We agreed on three rules: the AI's original value is never overwritten, every change is tied to a person, and changes to important fields need a short reason.

Q4 vendor invoices / Invoice_4471.pdf / History
Changed values (3)All fields (14)
FieldAI valueFinal valueChanged byReason
Phone number63% 555-555-555_ 555-555-5552 Priya S.14 Mar, 10:42 Last digit faint, checked scan
Invoice date57% 2 Jan 2023 1 Feb 2023 Priya S.14 Mar, 10:44 US date format
PO number96%, spot check PO-4471-8 PO-4471-B Rahul K.14 Mar, 15:10 B read as 8, checked against PO list
PO number was a high-confidence value changed at approval. Logged as a model miss and sent to the ML team.
Timeline
  1. UploadedAnil M., 14 Mar, 09:58
  2. Extracted by AI14 fields, model v2.3
  3. ReviewedPriya S. changed 2 values, confirmed 12
  4. Spot checkRahul K. fixed PO number at approval
  5. ApprovedRahul K., 14 Mar, 15:12
The AI value is struck through, not replaced, so anyone can see what the machine said and what a person decided. The model version is saved too, because the same document can be read differently after a model update.

This record also closes the loop with the ML team. A person changing an AI value is the clearest signal of where the model is weak, especially when the value was high confidence, like the PO number above.

11Outcome

It hasn't launched yet,
so here's how I'd measure it.

Five measures, using the research numbers as the starting point where we have them.

  • Time to review one document. Today it's about 15 minutes: 10 to type and 5 to check.
  • Documents per reviewer per day. Today it's about 30.
  • Fields still checked by hand. If people keep opening high-confidence fields, they don't trust the bands yet.
  • Documents waiting more than a day. Shows whether handoffs got better, not just the review screen.
  • How often people change an AI value. Especially high-confidence ones, since those are the misses the model doesn't know about.

What I learned

Start with the decision, not the AI.Work out what the person has to decide, then decide where the AI helps.

A score needs a reason.Confidence is only useful when it tells someone what to do next.

Design for the AI being wrong.The flow matters most on the documents where the model makes mistakes.

Less repeated checking, clearer decisions, and people stay in control.

Mohit Madan, Senior Product Designer, Bangalore  ·  mohitonline.com