Unicode vs InPage: Why Urdu Publishing Is Finally Moving On
The difference between Unicode Urdu and legacy InPage files, why InPage dominated Urdu publishing for decades, and how to convert old documents without losing your work.
For thirty years, InPage was Urdu publishing. Newspapers, books, government forms and academic journals were all set in it. Its files are still everywhere — and increasingly, they are a problem.
Understanding why explains a lot about the state of Urdu on the internet.
What InPage was, and why it won
InPage launched in 1994, at a time when no mainstream software could set Nastaliq properly. Word processors could barely handle right-to-left text, let alone the dense contextual ligatures Urdu calligraphy requires.
InPage solved this brilliantly for its era. It produced genuinely beautiful Nastaliq, handled the complex line-breaking Urdu needs, and gave Urdu publishers professional typesetting when nothing else could.
It became the standard because it was, for a long time, the only real option.
The catch: font encoding
To achieve this before Unicode was widely supported, InPage used font encoding. Urdu text was stored as ordinary Latin characters, and a specific font — typically Jameel Noori Nastaleeq — drew Urdu shapes in place of those Latin letters.
So the byte k in the file might display as ک, purely because of which font was applied.
Why this is a problem now
The text is not actually Urdu. It is Latin characters wearing an Urdu costume.
- Open the file without that exact font and you see meaningless Latin gibberish.
- Search does not work. Searching a database of InPage documents for کتاب finds nothing, because no such string exists in the file.
- Copy and paste breaks. Paste into any other application and the Urdu evaporates.
- Screen readers cannot read it. The accessibility implications are severe.
- Machine processing is impossible. No indexing, no translation, no text analysis.
What Unicode changed
Unicode assigns every Urdu character its own permanent identity. ک is U+06A9 in every file, on every system, in every font.
The consequences are substantial:
| InPage (font-encoded) | Unicode | |
|---|---|---|
| Portable across apps | No | Yes |
| Searchable | No | Yes |
| Works without a specific font | No | Yes |
| Copy and paste | Breaks | Works |
| Screen reader accessible | No | Yes |
| Web-ready | No | Yes |
Modern software — Word, browsers, phones, InPage's own recent versions — all handle Unicode Urdu natively. The technical reason InPage's approach existed no longer applies.
The migration problem
None of this helps with the archive. Decades of Urdu material sits in font-encoded files: newspaper back-catalogues, book manuscripts, government records, academic work.
Converting it is not trivial, for one specific reason.
There was never one InPage encoding
Different publishers, different InPage versions and different Nastaliq fonts used different mappings between Latin bytes and Urdu letters. There is no single authoritative table.
A converter can handle the common InPage 2.x / Jameel Noori encoding well, which covers a large share of real files. But an unusual publisher's internal font may map characters differently, and the conversion will produce plausible-looking wrong output rather than an obvious error.
This is why converted InPage text must always be proofread. It is a strong starting point, not a finished job.
How to convert
For text you can select and copy:
- Open the InPage document.
- Copy the text you need.
- Paste it into the Unicode Converter in InPage → Unicode mode.
- Read the output carefully against the original.
- Copy the corrected Unicode wherever you need it.
For scanned or image-based documents, conversion is an OCR problem rather than an encoding one, and Urdu OCR remains considerably less reliable than Latin OCR.
Related: text from PDFs
A neighbouring problem worth mentioning. Urdu copied out of a PDF often arrives as Arabic Presentation Forms — pre-shaped glyphs from the U+FB50–U+FEFF block rather than normal letters.
The symptoms look similar to encoding problems: text that appears fine but will not search and breaks when edited. The converter handles this too, in Presentation → Unicode mode.
What to do going forward
If you are producing new Urdu content, use Unicode. Every current tool supports it, including recent InPage versions. There is no longer any typographic advantage to font encoding — web browsers and Word render Nastaliq properly given a Nastaliq font.
If you maintain an archive, converting to Unicode makes it searchable and future-proof. Budget time for proofreading; the conversion is the fast part.
If you receive Urdu from others, normalise it once when it enters your workflow rather than working around problems repeatedly afterwards.
The wider point
InPage was not a mistake. It was the right solution to a real constraint, and Urdu publishing owes it a great deal.
But the constraint is gone. Urdu text that is genuinely text — searchable, portable, accessible, machine-readable — is what lets the language participate fully in the modern internet. That is worth the migration effort.
You can start with the free Unicode converter, which runs entirely in your browser.
Keep reading
Urdu Unicode Explained: Why Your Text Breaks and How to Fix It
Why Urdu that looks perfect in one app turns to squares in another — and how to repair it. A practical guide to Urdu Unicode, Arabic lookalike letters and Presentation Forms.
12 Urdu Typing Tips That Will Double Your Speed
Practical tips for typing Urdu faster and more accurately — spelling conventions, keyboard shortcuts, diacritics, and the mistakes that slow most people down.
The Best Urdu Keyboard in 2026: What Actually Matters
Not all Urdu keyboards are equal. Here is what separates a good one from a frustrating one — accuracy, typography, privacy and export — and how to judge for yourself.