Back to the blogThe Jyper blog / Model guides

Qwen, for everything you have to squint at.

Every DMC has the folder. The contract that only exists as a scan. The rate card someone photographed at a trade fair. The availability screenshot from WhatsApp. None of it is text until something reads it off the pixels, and reading it wrong means quoting a wrong rate.

By the Jyper team · August 23, 2026

Here is a quiet embarrassment of the AI world. The famous models, the ones on billboards, are mediocre at reading documents that arrive as pictures. Hand them a crisp PDF and they shine. Hand them a slightly crooked scan of a 2019 hotel contract with a stamp over the rates table and they start guessing.

The models that do not guess mostly come from one place. Chinese labs, whose models grew up reading dense documents in two writing systems, keep winning the reading tests. On the main public test for reading text out of images, Alibaba’s Qwen models sit ahead of Google’s Gemini and far ahead of GPT. And on the scoreboard where people vote blindly on image tasks, Qwen’s newest model is second in the world, the only non-Claude model near the top.

Best at: the squint folder

Think about what your files actually look like. The seasonal supplement table that lives inside a PDF as an image, so you cannot even search it. The photographed page where the fold hides half a column. The screenshot where the rate is in the image, the dates are in the caption, and the room category is implied by the thread above.

A rate you cannot search is a rate someone will eventually retype wrong.

Old-fashioned scanning software pulls the characters off the page and stops there. It does not understand that the table is a rate structure, that the footnote changes the cancellation terms, or that two columns describe two room categories. Qwen reads and understands in one pass: this is the high-season rate, this asterisk matters, this column is per person and that one is per room. That difference is exactly what the document-reading tests measure, and it is where Qwen keeps winning.

The catch: it reads better than it decides

We would not hand Qwen the delicate negotiation email or the tricky costing decision. On the scoreboards for judgment and long working sessions it sits mid-table, behind Claude and GPT-5.6. No shame in that. The specialist who reads every scan perfectly does not have to also be the person who decides the margin.

Speed, accuracy, cost

The three numbers. Accuracy: the best in the business at its specialty, reading text and tables out of images, per the tests above. Speed: unremarkable, around 50 words a second. Cost: mid-range, about $2 per million tokens read and $6 per million written, tokens being the small chunks of text AI is billed in, per Artificial Analysis. Notably cheaper than the flagship models that read scans worse.

That last sentence is the economics lesson of this whole guide in miniature. If you buy AI directly, your bill grows with usage, and the instinct to run everything through the most famous model means paying more per page for worse reading. The specialist is both better and cheaper at its own job. Matching the model to the task is not a nice-to-have; at scale it is the difference between an AI bill you notice and one you do not.

What it does in Jyper

In Jyper, Qwen works the squint folder. Scans, photos and screenshots go to it, and what comes out is structured data: rates with their seasons, supplements with their conditions, each traced back to the exact document it came from. Then Jyper’s own pricing engine does the arithmetic. The model reads; it never invents a number, and if it read something badly, the trace back to the source page is how a reviewer catches it in seconds.

The practical effect, and honestly our favourite thing about it: the squint folder stops being a folder of dread. It becomes just more data, as searchable and quotable as everything else you own.

Where the numbers come from: the OCRBench v2 public leaderboard (June 2026), the arena.ai image scoreboard (checked August 23, 2026), and Artificial Analysis for speed. Models change monthly. We re-check when new ones ship.
Questions people ask

What is the best AI for reading scanned contracts and photographed rate cards?

On the current public tests, Alibaba’s Qwen family. It leads the main image-reading benchmark ahead of Gemini and GPT, and ranks second on the blind-vote scoreboard for image tasks. That is why Jyper sends scans, photos and screenshots to Qwen.

Why not just use normal OCR scanning software?

Scanning software extracts characters and stops. It does not understand that a table is a seasonal rate structure or that a footnote changes cancellation terms. Qwen reads and interprets in one pass, which is the difference between a pile of text and data you can quote from.

Does Jyper let Qwen set prices?

No. Qwen reads documents into structured data, each item traced to its source page. Jyper’s own pricing engine computes every rate, margin and total. A model never invents a number, and the trace makes bad reads easy to catch in review.