We needed to replace the OCR engine without making users feel the migration
Free usage from the previous OCR vendor ended in September 2025. KYC still depended on identity-data extraction, while the capture flow, identity form, and surrounding logic were already running in production.
Vertex AI was already available in AstraPay's environment. The challenge was not only whether the model could read an ID correctly. The new output had to be stable enough to replace the old vendor response without forcing users to learn a new flow or making the app interpret model responses case by case.
I worked with engineering on the prompt, response contract, document detection, and cost tradeoffs. My focus was keeping the technical change behind the scenes while preserving a familiar experience in front of it.
The migration worked if users still saw a document photo turn into a form ready for review, while the system behind it could change OCR providers with less integration risk.
Three things still had to be true after the vendor changed
These constraints made the problem closer to system design than simply choosing the model with the best extraction result.
The new model sat behind a KYC flow users already knew
Users still captured a document and reviewed their identity data. The largest change happened in the extraction and validation layer before data reached the form.
User takes a photo of a KTP or SIM
The model reads the image and returns structured data
Strict JSON keeps fields and formats predictable
Extracted data maps into the existing form
User checks the fields and edits them when needed
Flow connections
- capture to vertex
- vertex to contract
- contract to prefill
- prefill to review
This diagram shows responsibilities across layers rather than backend implementation detail.
Output that reads well to a human can still be unsafe for an application
An LLM can return the right information in many different shapes. KYC needed a response that was much more boring and consistent.
Correct, but hard to predict
The model can add explanation, markdown, different field names, or other text the application does not need.
Every formatting variation becomes another integration edge case.Structured and backward-compatible
The response is constrained to strict JSON and stays aligned with the previous OCR schema. Missing information follows a defined behavior rather than being filled through model assumptions.
The identity form can keep using the mapping it already depends on.Try the OCR flowWe preserved the old contract and added only the information that could support product logic
The goal was not to make the response richer for its own sake. Each new field needed a reason to exist downstream.
The prompt acted like a specification for a product response
We defined which fields had to be returned, how their values should be formatted, and what should happen when information was not visible in the image. The model was instructed not to invent data it could not read.
I iterated on the prompt with engineering until the response was consistent enough to map directly into the identity form. In this context, prompt quality was less about how complete the answer sounded and more about how little extra interpretation the app needed.
documentType then created a foundation for handling more than one identity document. We treated matchRate more cautiously as an additional signal rather than a number that could automatically decide whether a user's data was right or wrong.
For product integration, predictable output is more useful than output that sounds intelligent.
See how the new OCR still fits the same KYC experience
This walkthrough shows document capture, prefill, and correction from the user's perspective. The screens are illustrative, while document routing and correction behavior follow the product logic from this project.
The structured response also created a path beyond KTP
documentType let one OCR layer tell the app what kind of document it had received before data moved further through the flow.
KTP was the primary capture path
The earlier integration followed a contract built around the main identity flow and did not return document type as part of the OCR response.
Adding another document required more handling around the old contract.KTP and SIM could share the same foundation
The new response identifies the document type. SIM could then be added as an alternative without creating a completely separate capture experience.
The capability changed behind a flow that remained familiar to users.A cheaper model does not help if users have to correct more of the form
Vertex AI introduced an inference cost for each request. We looked at token usage, model pricing, and expected KYC volume so the model choice would remain reasonable at production scale.
One scenario came to roughly USD 0.00102135 per request, about Rp16 at the time. We used that figure to understand the order of magnitude, not as a claim that every request would always cost exactly the same.
Cost still had to be read alongside prefill quality. Saving a few rupiah per request would not mean much if correction increased and users lost the main benefit of OCR.
The OCR layer changed while the user's job became slightly lighter
The evidence available to me covers correction, document capability, and a planning cost scenario. I keep those as separate kinds of evidence rather than treating them as one success metric.
I kept the 10% estimate as an estimate instead of turning it into a more precise-looking chart
The project records an approximately 10% reduction in user edit rate, but the absolute before-and-after edit-rate percentages are not available in the portfolio material. I therefore keep it as a relative estimate rather than inventing a baseline that looks directly measured.
The same boundary applies to cost. Rp16 represents one model scenario from planning, not a fixed production cost for every request.
Field-level correction would be more useful than one aggregate edit rate
The next iteration should separate correction rate by field and document type. That would show whether errors cluster around names, dates, addresses, or a particular document.
I would also compare matchRate with the corrections users actually make. If the relationship is consistent enough, the signal could then be considered for more selective product treatment.
