DocsAPI LogoDocsAPI

OCR for Healthcare: Medical Documents, Accuracy, HIPAA

Nupura Ughade
Nupura Ughade
|
October 5, 2026
|
14 min read
In short

OCR can turn scanned or photographed medical records, insurance cards, prescriptions, lab reports and claims into usable fields. Clean print reads well, but handwriting, faxes and medical codes need checking. Software is not HIPAA compliant by itself: the vendor agreement (BAA) and controls decide.

A fax lands at 4:50 pm, a little crooked, with a drug name handwritten in the margin. Somebody has to decide whether it says hydroxyzine or hydralazine before it goes anywhere near a chart or a claim. OCR reads most of that page well. The parts it cannot read safely are exactly the ones that matter, and this guide is about telling the two apart. It is not legal advice.

What can OCR do with medical documents?

OCR turns a scan, fax or photo of a medical document into text a computer can search. Extraction software then picks out fields such as member ID, drug, test result and diagnosis code, and checks them against code lists. OCR does not decide what a record means or whether a claim gets paid.

Most failures happen after the reading step. Software must decide which text is the member ID and which is the group number, then validate values against code lists, check digits and arithmetic, and send doubtful fields to a person. Here is how that plays out for the five document families people search for most.

  • OCR for medical records turns charts and outside records into searchable text and splits multi-report PDFs.
  • OCR for insurance cards reads member ID, group number, payer ID and pharmacy numbers from both sides.
  • OCR for prescriptions reads drug, strength, quantity, directions and prescriber identifiers.
  • OCR for lab reports reads tests, results, units and reference ranges, then maps tests to LOINC codes.
  • OCR for insurance claims reads claim forms such as the CMS-1500, plus EOBs and remittance advice.

A naming trap: in HIPAA writing, "OCR" usually means the HHS Office for Civil Rights, which enforces the HIPAA rules. This page is about optical character recognition, the software. The two meet in the HIPAA section below.

Medical document cheat sheet: fields, codes and the failure to watch

Eight document types, what OCR pulls from each, the codes and identifiers involved, and how each goes wrong. We checked the code details on September 21, 2026 against CMS, the FDA, the NUCC (which maintains the CMS-1500), X12, the CDC ICD-10-CM guidelines, NCPDP, the AMA and LOINC.

DocumentWhat OCR extractsCodes and identifiersThe specific failureRead more
Insurance card (both sides) Member ID, group number, plan name, payer ID, pharmacy numbers, claims address Payer ID, RxBIN (6 digits), RxPCN, RxGRP Plan name captured instead of payer ID. A front-only photo misses the pharmacy numbers on the back Insurance card OCR, eligibility verification
Prescription Drug, strength, quantity, directions, prescriber, NPI, DEA number NDC (10 digits on the label, 11 for billing), NPI (10 digits), DEA number (2 letters, 7 digits) A zero added in the wrong NDC segment. A misread NPI or DEA digit that nobody checks NDC extraction, NPI verification, DEA verification
Lab report Test, result, unit, reference range, specimen, date LOINC (names the test, not the result), local lab codes A bare "Glucose" hides the specimen. LOINC 2345-7 is serum or plasma, 2339-0 is whole blood LOINC code extraction
Medical record or referral Diagnoses, medications, dates, providers ICD-10-CM (3 to 7 characters), CPT (5). C-CDA or FHIR if structured A code missing its 7th character is invalid. A valid C-CDA file can hold an empty list Medical coding, C-CDA, FHIR
Claim form (CMS-1500, UB-04) Patient, insured, up to 12 diagnoses (Item 21), service lines (Item 24) ICD-10-CM without the decimal point in Item 21, CPT or HCPCS, pointer letters A to L in 24E, NPI A pointer letter mistaken for a modifier in the crowded Item 24 grid CMS-1500 processing
Explanation of benefits Billed, allowed, plan paid, deductible, copay, coinsurance No standard layout or codes Right number under the wrong label. Check the arithmetic EOB parsing
Remittance or denial Paid amount, adjustments, group code, reason codes CARC, RARC, group codes such as CO and PR CARC 16 read without its remark code. RARC MA130 means no appeal rights, so refile 835 remittance, denial codes
Prior authorization or referral fax Patient, provider, CPT, units, dates, authorization number CPT, ICD-10-CM, NPI Standard fax is about 204 by 98 dots per inch and black or white only, so fine print and handwriting are lost Prior authorization, medical fax

How accurate is OCR on medical documents?

There is no single number. In our benchmark, clean printed English scored 97 to 99% on every engine, but that benchmark did not include medical documents. Handwriting and phone photos scored much lower, and a correct read can still be a wrong code. Measure accuracy on your own worst documents before you trust any headline figure.

What our benchmark measured, and what it did not

We tested eight engines on 1,900 real documents, including invoices, bank statements, receipts, handwritten forms and academic papers, but no medical documents. These rows show what clean print, handwriting and phone photos do. DocsAPI is our product. Method: OCR accuracy benchmark.

Document categoryTesseract 5.xPaddleOCRDocsAPI
Typed English (clean)97-99%97-99%97-99%
Handwritten forms (mixed neatness)61%73%78%
Phone-photographed receipts58%82%76%

Handwritten forms are not handwritten prescriptions, so read this as direction, not a forecast. Going from clean print to handwritten forms cost every engine 19 to 38 points (97-99% minus 78%, 73% or 61%).

What published medical studies found

  • Mayo Clinic, JAMIA Open, 2026. OCR plus a large language model processed 1,303 scanned outside-record PDFs from 116 institutions (breast cancer care). F1 scores (1.0 is perfect) were 0.95 for splitting multi-report PDFs, 0.96 for classifying documents and 0.90 for extracting study dates. In a pilot of 45 records, clinicians reported 2 classification errors and 1 date error, and estimated review time fell 40%.
  • PLoS One, 2024. 38 volunteers recorded medication lists from photos of pharmacy notebooks, by OCR and by typing on a smartphone. Errors (characters omitted or misread) were 0.62% for OCR (30 of 4,814) and 1.10% for typing (53 of 4,814). A six-medication list took 18 seconds by OCR and 144 by typing, 8 times faster (144 ÷ 18 = 8). It used only six photos in a simulation, and several authors were paid by the OCR tool's maker.

The useful comparison is against typing, not perfection. Both studies are small and narrow, so neither predicts your mix of documents. Character accuracy can also look excellent while a claim bounces: a correct read of "Glucose" still needs the right specimen code. Our post on why 99% accuracy can still mean wrong data explains the gap, and how to measure OCR accuracy shows how to score your own documents.

How do you catch OCR errors before they reach a claim?

In healthcare a misread is not a typo. It can be the wrong drug, the wrong patient or a denied claim. Use checks that do not depend on the OCR engine: check digits for identifiers, official code lists, arithmetic on payment documents, and a person for low-confidence fields. A wrong value that looks right is worse than a blank field, so send doubts to review instead of guessing.

  1. Identifiers. NPI and DEA numbers carry check digits. A failed check catches a misread digit. A pass does not prove the person exists, can prescribe or can bill.
  2. Drug codes. Convert a 10-digit NDC to the 11-digit 5-4-2 billing form by adding a zero to the short segment, not to the front of the string.
  3. Diagnosis codes. An ICD-10-CM code has 3 to 7 characters and is invalid without a required 7th. CMS says new codes take effect October 1, 2026, so check the list for the date of service.
  4. Payment arithmetic. On an EOB, allowed amount minus plan payment should equal what the patient owes.
  5. Reason codes in pairs. X12 says CARC 16 requires at least one remark code.
  6. Fields, not pages. Score each field, and score 50 of your messiest real documents against hand-typed answers before you scale.

Here is a check for items 1 and 2, in plain Python 3. The NPI rule and its example (1234567893) come from the CMS check digit document. The NDC rule, adding a leading zero to the appropriate segment, comes from the FDA's March 2026 final rule text.

def npi_ok(npi):
    if len(npi) != 10 or not npi.isdigit():
        return False
    total = 0
    for i, ch in enumerate(reversed("80840" + npi)):
        d = int(ch)
        if i % 2 == 1:
            d *= 2
            if d > 9:
                d -= 9
        total += d
    return total % 10 == 0

def ndc_to_11(ndc):
    labeler, product, package = ndc.split("-")
    return labeler.zfill(5) + "-" + product.zfill(4) + "-" + package.zfill(2)

We ran it. Output:

npi_ok("1234567893")  True
npi_ok("1234567890")  False
0574-4072-05  to  00574-4072-05
51655-089-52  to  51655-0089-52
58118-1397-3  to  58118-1397-03
0 of 90 single-digit misreads pass

The last line tries every one-digit misread of that NPI (10 positions times 9 other digits). None pass. The NDC strings only show the pattern, so confirm real labeler codes in the FDA directory. See NPI verification and DEA number verification for what these checks cannot prove.

Is OCR for healthcare HIPAA safe?

OCR software is not HIPAA compliant or non-compliant by itself. HIPAA applies to covered entities and their business associates, and HHS does not endorse or certify products. What matters is the signed business associate agreement (BAA), the vendor's safeguards, how long files are kept, and what is done with them. This is not legal advice.

Every document on this page is protected health information (PHI) once it ties a patient's identity to health care or payment. HHS's examples are a medical record, a laboratory report and a hospital bill. HHS's Office for Civil Rights (the other OCR) says it "does not endorse, certify, or recommend specific technology or products," so ask what a "HIPAA compliant" badge rests on. Five things decide the answer.

  1. The BAA. A vendor that creates, receives, maintains or transmits PHI for you is a business associate, even if it stores only encrypted data and holds no key. HHS says using a cloud provider to maintain ePHI without a BAA is a violation. Under 45 CFR 164.504(e) the contract must limit use of the data, require safeguards and, where feasible, return or destroy PHI at the end.
  2. Access controls and logs. The Security Rule requires administrative, physical and technical safeguards, including access control, audit controls (records of who touched what), authentication and transmission security. Ask who at the vendor can open your documents and how that is logged.
  3. Retention and deletion. HIPAA does not set how long medical records must be kept. HHS says state law generally does. Ask how long the vendor keeps a file after processing, whether you can delete on demand, and whether your documents train models. A business associate may use PHI only as its contract permits.
  4. De-identification. HHS recognizes Expert Determination and Safe Harbor, which removes 18 listed identifiers. De-identified data is no longer PHI, but redaction that misses a name in a margin note has not de-identified anything. A business associate may de-identify for you only if the BAA authorizes it. See HIPAA de-identification automation.
  5. Minimum necessary. Send and keep only what the job needs. It is a different standard from de-identification and has exceptions, such as treatment disclosures and signed authorizations. See HIPAA minimum necessary redaction.

Running an open source engine yourself, such as Tesseract or PaddleOCR, means no outside company holds the data. The Security Rule duties, including a risk assessment of your ePHI, are then entirely yours.

What do hospitals, pharmacies and claims adjusters digitize?

They digitize whatever still arrives on paper, fax or photo. Hospitals handle referrals, outside records, lab results and claim forms. Pharmacies handle prescriptions, insurance cards and drug labels. Claims adjusters handle the claim file: forms, photos, estimates, reports and bills. Each group reads different fields, so OCR is set up per document type, not per industry.

OCR for hospitals

In a November 2019 Medical Group Management Association poll (1,581 responses), 89% of healthcare leaders said their organization still used a fax machine, citing record sharing, referrals, lab results, payer communication and pharmacy communication. The poll is old, but it shows which documents keep arriving as images. Hospital teams also scan release of information requests (release of information processing), Medicare Advance Beneficiary Notices (ABN processing), and the chart documentation behind denials and risk scores (medical necessity documentation, HCC risk adjustment coding).

OCR for pharmacies

Prescriptions arrive electronically or as paper, fax or photo. For Part D controlled substances, CMS counts a prescriber as compliant at a 70% electronic rate and notes that e-prescribing reduces pharmacy calls to clarify written prescriptions, so paper does not vanish by rule. The pharmacy-side identifiers are the NDC, the prescriber's NPI and DEA number, and the RxBIN, RxPCN and RxGRP on the card. Pharmacy claims use NCPDP, separate from medical claims, so scan both sides of the card.

OCR for claims adjusters

The NAIC handbook defines an adjuster as a person who investigates claims, determines coverage, examines relevant documents and inspects property damage. Some states limit adjusters to a specialty such as auto, homeowner or workers' compensation. NAIC consumer guides list what fills the file: photos, damage lists, receipts, repair bids, police report numbers and proof of insurance. Where a claim involves an injury, the file can also hold the medical documents above. OCR speeds the reading. Coverage stays with the adjuster and the policy.

How do insurance companies process claims documents?

Insurers check who is covered, review the claim and its documents, decide what to pay, and send a written explanation. Health claims follow HIPAA standard transactions: an 837 claim goes in and an 835 remittance comes back. Paper forms are scanned with OCR. Auto and property claims add an adjuster who inspects the loss and examines the documents.

Health insurance, step by step

  1. The claim arrives. Providers send an ASC X12 837 (professional, institutional, dental) or NCPDP (retail pharmacy) claim, usually through a clearinghouse. Some send paper: the CMS-1500 or UB-04. CMS says Medicare takes the CMS-1500 from providers with a waiver from the electronic requirement, most paper claims to carriers are scanned with OCR, and photocopies are not accepted by all carriers because they cannot be scanned.
  2. Checks before payment. The payer verifies eligibility (270/271), and some services need prior authorization. See eligibility verification automation and prior authorization automation.
  3. Adjudication. This is the payer's decision process. Plan rules and coverage policies are applied to each service line coded with the HIPAA standard code sets (ICD-10-CM, CPT and HCPCS, NDC). Each line can be paid, reduced or denied.
  4. The explanation. The provider gets an 835 remittance advice with CARC and RARC codes. The patient gets an EOB or, in Original Medicare, a Medicare Summary Notice. Private insurers called Medicare Administrative Contractors process Medicare claims.

Auto and property insurance, and where OCR fits

The NAIC's guides describe a shorter path: report the loss, the insurer assigns an adjuster, and the adjuster assesses the damage to set the settlement amount. OCR sits at intake, reading scanned claim forms, attachments and remittances into data. The payment decision stays with the payer's rules and its people. Our post on revenue cycle management document intake explains why slow intake delays billing and shows a 20-transaction check for timing your own.

Where to begin

Start with the documents where a misread hurts least, and earn trust before you go near prescriptions.

  • Clinic or hospital front desk: start with insurance cards and faxed referrals. Score 50 real ones and read the payer ID and authorization numbers yourself.
  • Pharmacy: start with the prescription and both sides of the card. Run the NPI, DEA and NDC checks and send doubtful drug fields to a pharmacist.
  • Billing team: start with remittances and denials. Read CARC and RARC as a pair (denial code extraction).
  • Adjuster or insurer: start with the documents that fill the file, and keep the coverage decision with people.
  • Anyone handling patient data: get the BAA first, and test with synthetic or de-identified samples.

To test a document API, DocsAPI (our product) lists prescriptions, lab reports and claims on its medical documents page. Its published policy says HIPAA-eligible with a BAA on request, SOC 2 Type II, and content deleted after processing and not used for training. Confirm that in the contract, as with any vendor. For contracts or IDs instead, see OCR for legal documents, OCR for KYC verification or OCR for passports, ID cards and driver's licenses.

Sources

We opened each source September 20 to 21, 2026. hhs.gov blocks automated fetching, so we read archived 2024 and 2025 copies of the HHS pages, which may have changed. The payer ID description relies on our insurance card OCR post. Our benchmark has no medical documents, and we found no published head-to-head engine test on them. Nothing here is legal, medical or coding advice.

Common questions

Frequently asked questions

In HIPAA writing, OCR almost always means the HHS Office for Civil Rights, the agency that enforces the HIPAA Privacy, Security and Breach Notification Rules. It has nothing to do with optical character recognition software, although the two meet when a hospital scans records. Search results for this topic mix both meanings, so check which one a page is using.

We could not verify a published head-to-head test on medical documents, and our own benchmark did not include them. In that benchmark the winner changed by document type, so no engine led everywhere. Pick two or three candidates, run 50 of your own worst documents through them, and score the results against hand-typed answers.

Sometimes, with more errors than printed text. In our benchmark, handwritten forms scored 61% to 78% against 97% to 99% for clean print, and forms are not prescriptions. CMS notes that e-prescribing reduces pharmacy calls to clarify written prescriptions. Send any doubtful drug, strength or quantity to a pharmacist rather than accepting it from OCR.

We did not test scanners, so we cannot name one. The card layout gives one rule: capture both sides, because pharmacy routing numbers and the claims address usually sit on the back. Scan a few real cards and read the small print on the result, especially the payer ID, before you buy. A fax-quality image is too coarse for small print.

Only if the provider will sign a business associate agreement. HHS says a cloud service that processes or stores ePHI for you is a business associate, and that using one without a BAA violates the HIPAA Rules. If a site will not sign one, do not upload documents with patient identifiers. This is not legal advice.

A business associate may use PHI only for what its contract with you permits, and HHS says a business associate may de-identify data for you only if the agreement authorizes it. So the answer is in the BAA. Ask for the use and retention terms in writing, and get a plain yes or no on model training.

No. A clinical note rarely contains the code itself. It contains a description that someone must match to the best ICD-10-CM or CPT code under the coding guidelines, which is a classification decision, not a reading task. OCR gets the words into the system. Our medical coding automation post explains why accuracy stalls when coding is treated as extraction.

An EOB goes to the patient and explains what was billed, what the plan covered and what the patient may owe, in a layout each payer chooses. A remittance advice goes to the provider, usually as a standard 835 file with coded adjustment reasons. They describe the same claim for different readers.

It depends on the field. A wrong payer ID typically gets the claim rejected by the clearinghouse before a person sees it, with a reason that does not point back to intake. A wrongly padded drug code can produce an invalid NDC denial that looks like a coverage problem. Both are found late, which is why the checks belong at capture.

If a document links a patient's identity to health care, payment for care or a health condition, yes. HHS gives a medical record, a laboratory report and a hospital bill as examples, because each carries the patient's name or other identifying information. An insurance card in a provider's patient file is tied to payment for care, so treat it as PHI too, and ask your privacy officer if unsure.

Typed medical records, lab reports and printed prescriptions read well. Handwritten prescriptions and clinician notes are much harder, and a misread drug name or dose is a patient-safety risk, so medical record OCR should send low-confidence text to a person and never fill in orders automatically.

Yes. Commercial OCR SDKs and open-source engines such as Tesseract and PaddleOCR can run inside your own network, which helps keep protected health information in-house. A medical document OCR SDK still needs your own checks for codes such as ICD-10 and NDC, and a HIPAA risk assessment covering where images and text are stored.

Nupura Ughade

Content Marketing Lead, DocsAPI

Nupura Ughade creates clear, insightful content on OCR, document AI, and fintech. She combines technical depth with real-world finance use cases to help engineers and operations leaders navigate digital transformation with confidence.

Want to see it on your own documents?

Try our free OCR tool in your browser, or book a demo to see how DocsAPI reads the document types covered in this guide.