Most of the websites I build solve a business problem: get more enquiries, rank higher, convert more visitors. Farouq on Cinema is a different kind of project. It's a bilingual digital archive of Egyptian writer, director and producer Farouq Abdulaziz's film criticism — magazine articles published across the Arab world from 1968 to 2005, his long-running show Cine Club, and decades of television interviews. The goal wasn't conversions. It was making sure real, significant cultural material didn't just sit in physical magazine pages nobody could search.

1. Digitizing print material means deciding what "the source" actually is

OCR (optical character recognition) turns a scanned magazine page into searchable text, but OCR is never perfect — old print, faded ink, and non-standard layouts all introduce errors. The decision that mattered most on this project wasn't which OCR tool to use, it was keeping the original scanned image alongside the extracted text for every single article, not just the text alone. If a reader — or a researcher citing the piece — wants to verify a word against the real printed page, they can, instead of trusting a machine transcription on faith.

2. Metadata is what turns a pile of scans into an archive

A folder of scanned images isn't an archive; it's a folder of images. What makes it searchable and citable is structured metadata attached to every piece — in this case, the publication date and the specific magazine each article ran in, stored alongside the text and the image. That's what lets a visitor filter to "everything from 1975" or find a specific magazine's coverage, rather than scrolling through decades of undifferentiated content.

3. Video and print need to live in the same place, not separate silos

Farouq's interviews — on YouTube, and on television including Al Jazeera and Kuwait TV — are as much a part of the record as his written criticism. Embedding that video content alongside the article archive, rather than leaving it scattered across whatever platform originally hosted it, means a visitor researching his work doesn't have to piece together his career from five different sources. The site includes a dedicated section for Cine Club, his best-known show, for the same reason — it's a distinct enough body of work to deserve its own place rather than being folded into the general article list.

4. Bilingual isn't optional when the source material is Arabic

The overwhelming majority of the original articles are in Arabic, published in Arabic-language magazines for an Arabic-reading audience — so an English-only site would misrepresent the material from the start, and a machine-translated layer bolted on top would do the same work badly. Building the site genuinely bilingual, with proper right-to-left layout and separate content structure per language rather than an auto-translate widget, was a requirement of the material itself, not a nice-to-have. This is the same distinction covered generally in SEO services in Kuwait's discussion of Arabic search, and in more depth for a commercial context in bilingual Arabic/English SEO in Kuwait.

5. SEO for an archive means one thing above all: per-article discoverability

An archive with beautiful design and broken search is worse than useless — it hides exactly the thing people came for. Every article needed its own clean, stable URL and its own structured data identifying it as a distinct piece of content with a real publication date, so a specific article can actually surface in a Google search rather than being invisible inside one long undifferentiated page. Technical SEO usually gets framed around commercial goals — rankings, traffic, conversions — but the underlying discipline (clean URLs, correct markup, genuine content structure) is exactly what a cultural archive needs too, just aimed at "can a researcher find this specific 1982 article" instead of "can a customer find this product."

Frequently asked questions

Is OCR accurate enough to trust for historical print archives?

Not on its own — OCR accuracy on decades-old print varies with scan quality and layout, and errors are inevitable at scale. That's exactly why keeping the original scanned image next to the extracted text matters: the text makes the archive searchable, and the image is the actual source of truth a reader can fall back on.

Does a bilingual archive site need separate URLs per language, or one page with a toggle?

Separate, language-specific pages generally perform better for SEO than a single toggling page, since each language version can be indexed and ranked independently in its own language's search results.

What kind of project is a digital archive best suited for?

Any body of work at real risk of being lost or simply unfindable — a writer's or journalist's back catalog, a family or institutional archive, historical records tied to a specific person or organization. The common thread is print or broadcast material with genuine, lasting value that currently has no searchable digital home.

The takeaway

A digital archive isn't a smaller version of a business website — it's a different kind of project with a different goal: preservation and discoverability of material that already has value, rather than persuading a visitor to buy or book. The same technical discipline (clean structure, real metadata, genuine bilingual support, proper SEO) still applies, just pointed at a different outcome.

I build websites for businesses and for projects like this one — archives, historical records, and specialist content that need real technical care. If you have material that deserves a proper digital home, feel free to get in touch.

Related reading

Adnan Basra

Written by Adnan Basra

Senior Web Developer & E-Commerce Manager based in Kuwait, with 13+ years building websites and driving organic growth for businesses. Get in touch →

Have an archive or specialist project in mind?

Preservation, searchability and bilingual support — tell me what you're working with.

Ask on WhatsApp Contact Me