跳到主要內容
AI News HubLIVE
來源內容 · 翻譯待補全6 分鐘閱讀

待翻譯:How to Parse a Form | Unstructured

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Featured File TransformationLLM How to Parse a Form Oct 5, 2026 Featured File TransformationLLM How to Parse a Form Oct 5, 2026 Authors Ajay Krishnan Dev Rel Engineer, Unstructured In this article Join our newsletter to…

待翻譯:How to Parse a Form | Unstructured
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Featured File TransformationLLM How to Parse a Form Oct 5, 2026 Featured File TransformationLLM How to Parse a Form Oct 5, 2026 Authors Ajay Krishnan Dev Rel Engineer, Unstructured In this article Join our newsletter to receive updates about our features. In this article Forms are everywhere in the parts of a business that deal with real decisions. Think of a grant application, an insurance claim, or a loan file. Forms like these are how information gets into a company, and usually how a decision about money, eligibility, or a person gets made. A form also looks like the easiest kind of document to read. Someone fills one out in a couple of minutes, and you can glance at it and take in every answer without really thinking about it. So it seems fair to expect a computer to do the same, instantly, across thousands of them at once. But, it is actually one of the harder documents to read automatically, and what makes it look easy is the same thing that makes it hard. A person reads a crooked scan, a box ticked in pen, and a slightly different layout on every page, without noticing any of it. An automated reader does not get any of that for free. Most frontier models can read a form. The problem now is that forms are not really text, and treating them like text can cause problems. The tricky part is that nothing obviously breaks. The output still looks fine. You only find it later when the value you needed is sitting under the wrong label or gone, across a few thousand documents. A form is not a wall of text. A lot of the information in a form comes from the layout, not just the text. A checkbox is a yes or a no. A blank field still means something even when it is empty. A label owns the box sitting next to it, and only that box. A total on line 9 is the sum of the lines above it. None of that comes from the words alone. It comes from how everything is arranged on the page. If you read the form like a normal paragraph, you keep the words but lose the layout. So, what goes wrong when we parse a form? There are a few common problems. The value ends up under the wrong label. Most OCR techniques and VLMs try to read a page in the usual order: left to right and top to bottom. That works pretty well for normal text. Forms are different, because the closest piece of text is not always the field that value belongs to. A value in one column gets read into the field beside it, and now the county is in the address line and the county field is empty. The model got all the words right, but it still got the answer wrong. The same permit form. In the source, “Kent” sits in the County column. Unstructured keeps County = Kent. Read as flat text by an VLM, “Kent” lands in the address field, County comes back empty, a “(cont.)” and texts are hallucinated. Read in order, a header reads straight across the page: “City of Grand Rapids” while a VLM interpreted it as “CITY GRAND OF … RAPIDS.” This can also cause the model to invent structure that isn't actually there. When the reading order gets confused, a model will sometimes decide a section has started over and add a heading like "(continued)" that appears nowhere on the page, so now you have the wrong value under the wrong field, along with text that was never on the original form. The page has marks that are not text. Scanned forms also have a lot of things that can look like text but aren't actually text. A corner registration bracket. A punch hole. A stray pen mark. A redaction bar. If the model tries to read everything as text, those marks can end up in the output. The corner registration bracket on the source form gets transcribed as a stray character in the flat-text output. There is another problem too: invented fillers. When a field is empty, VLMs sometimes add their own placeholder to represent it, so now the output contains words the document never contained and you have to strip them back out before anyone downstream can trust the text. So how often does this actually happen? Everything above is a single observed case, one form where a value slid into the wrong column or a mark turned into a stray character. Each one is real, but on its own seems alright. To understand how often these errors actually happen, we ran a controlled test. We created 40 small business grant applications from the same one-page form: 13 clean digital PDFs, 13 ordinary scans, and 14 rough scans with tilt, low resolution, noise, faint pen marks, and cramped handwriting. Every form had a known answer key. We ran each one through Unstructured Transform and through a frontier model, GPT-5.6 Sol, giving the model the page image and asking it for the same structured output. Then we scored two things: how many individual fields each system read correctly, and how many whole forms came back with no errors at all. Field accuracy hides the number that matters. Field accuracy tells you how many individual fields a system reads correctly across all the forms. Each business name, EIN, dollar amount, or checkbox counts as a field. If a system gets 99 out of 100 fields right, its field accuracy is 99%. Both systems scored high here: 99.5% for Transform and about 98.7% for GPT-5.6 Sol. A single point apart, which sounds like a tie, and that is exactly the trap. A form can have almost every field right and still need a correction, so the more useful measure is how often every field on a form is correct. Unstructured TransformGPT-5.6 Sol Field accuracy 99.5% 98.7% Forms with zero errors 34 / 40 (85%) ~20 / 40 (51%) Digital forms with zero errors 13 / 13 8 / 13 Ordinary scans with zero errors 13 / 13 8 / 13 Forms with an incorrect name, EIN, address, or phone 0 12–18 Forms with a wrong dollar amount 1 1 Transform returned 34 of the 40 forms with no errors, or 85%. The frontier model averaged about 20 of 40, or 51%. That is the number that tells you how many forms needed a human: 6 out of 40 for Transform, roughly 20 for the VLM. A hard page where both models messed up. Transform made no errors on the 13 digital forms or the 13 ordinary scans, and all of its misses were on rough scans. GPT-5.6 Sol made errors in all three groups, including a business name on a clean digital PDF with nothing hard about it: GPT-5.6 Sol read “Ledger & Line CPAs” as “Ledger & Lime CPAs.” One letter changes the name on the record. Across the 40 forms, GPT-5.6 Sol got a name, EIN, address, or phone number wrong on 12 to 18 forms, while Transform got those fields right on all 40. What this looks like in practice. The errors are easy to miss, because the extracted text still looks like a real answer unless a person is checking that field against the source. That turns into real work in two places: Human review. About half the forms the VLM processed needed a correction, compared with 15% for Transform. Across 10,000 forms, that is roughly 4,900 to check and fix with the VLM against 1,500 with Transform. Someone has to do each of those by hand. Customer records. The VLM got a name, EIN, address, or phone wrong on 12 to 18 of the 40 forms. Transform got all of them right on all 40. A wrong value that looks plausible is the kind that survives review and lands in your records. And it costs a fraction as much! Transform was also cheaper in this test. Its best profile cost $0.015 per page against about $0.113 for GPT-5.6 Sol, using the promotional pricing applied during the test. SystemUnstructured Transform (best)GPT-5.6 Sol Per Page $0.015 $0.113 40 forms $0.60 $4.53 10,000 Forms $150 $1,130 1,000,000 Forms $15,000 $113,000 That works out to about 7.5 times the cost per page for the frontier model. At its standard pricing the gap is closer to 11 times. The larger-volume figures are projections from those per-page rates, and both systems took around 40 seconds per page. So the two costs point the same way. The frontier model charges several times more per page, and it hands back several times more forms for a person to fix, which means the cheaper system is also the one that leaves you less manual work. How does Unstructured process forms? Here at Unstructured, we take a multi-agent approach. We break a page into its parts and use a different model for each. Every part of a page that holds information is found first, then handed to the model best suited to it, which pulls the text out and formats it for what that block actually is. So what you get back is neatly formatted JSON with every element from your form, and each element carries the coordinates to where it sat in the original document. KeyValue 1 Legal name of businessBlue Heron Bakery LLC 3 Employer identification number (EIN)84-2917365 4a Street address (incl. suite)1180 Larkspur Ave, Suite 4 5 Entity type (check one)☐ Sole proprietor ☑ LLC ☐ S corporation ☐ C corporation ☐ Partnership ☐ Other Using the output. Since the output is already structured, using it is pretty simple. You don't need to define a schema first or spend time rebuilding the structure afterward. # One-time setup: # pip install unstructured-transform-client # export UNSTRUCTURED_API_KEY="your-key-here" import os from unstructured_transform_client import TransformClient client = TransformClient( api_key=os.environ["UNSTRUCTURED_API_KEY"], server_url="https://transform.unstructured.io", ) with open("filename.pdf", "rb") as f: result = client.parse.run( input=f, output="elements", # "elements" or "markdown" profile="best", # "balanced" is the everyday default; "best" favors quality ) for element in result.elements: print(element) Each form element comes back with its key-value table as HTML, so you can drop it straight into whatever consumes tables, a dataframe, a database, or the next step of an agent, without writing a second extraction pass to rebuild the pairs you would otherwise have lost. You can see where every value came from. Every element also comes with its coordinates on the page, so you can see exactly where a value came from. When a number matters and a human has to check it, or an agent has to act on it, being able to trace a value back to its box on the page is the difference between output you can audit and output you have to hope is right. Bring us your hardest forms! Forms can have a lot of small errors that are hard to notice, so this is one of the areas we focused a lot on. If you have a form that doesn't work well with your current parser, we'd love to try it. Test out your own forms at transform.unstructured.io! Join our newsletter to receive updates about our features. Related Articles Use Case Use Case: Consumer Goods Industry Jun 7, 2025 Unstructured Fine-tuning How We Taught an AI Agent to Fix Our Training Data Apr 23, 2026 Ajay Krishnan Use Case Use Case: AI Course of Action Generation and Analysis Dec 28, 2024 Unstructured

展開要點與分析

文章情報

工程師中級

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Featured File TransformationLLM How to Parse a Form Oct 5, 2026 Featured File TransformationLLM How to Parse a Form Oct 5, 2026 Authors Ajay Krishnan Dev Rel Engineer, Unstructure…

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。