News and media

Searching a news library by what happened, not what it was called

Keyword search fails because nobody named the file for the question you are asking now. DeepFile indexes what is in the material — text, images and video.

Ankit Goel 3 min read

A news organisation’s library is the most valuable thing it owns and the hardest thing in the building to use. Not because the material is missing — it is all there, decades of it — but because the only way most systems can look is by name, and nobody named a file for the question you are asking today. Context search indexes what is actually in the material, so you ask for it the way you would describe it to a colleague.

Why keyword search stops working at scale

Keyword search works when you already know the words in the thing you want. In a small archive that is often true. In a large one it almost never is. The photograph you need was filed as a number. The interview was named after the date, not the subject. The rights letter is in a folder called “misc”. The video was never labelled at all, because labelling video by hand is a job nobody finishes.

And the question changes. The story that made a clip worth keeping in 2011 is not the story that makes it useful now. A library organised around the old question cannot answer the new one.

What “by context” means

We built DeepFile for this. It is an agent rather than a search box: it indexes what is in your library, finds what you describe, and then answers questions about what it found. Indexing by context means it reads the material and builds an index of what each item is about, rather than what it was called. For documents and scans, that is the text inside them, including text recognised inside photographs. For images, it is what is pictured. For video, it is the people in it and what happens in it — footage becomes searchable by who is on screen and by describing the scene, with people and crowds tracked through the shot so a result is a stretch of footage with a beginning and an end, not a still.

All three sit in one index. That is the part that matters for a newsroom, because a real question crosses all three: the minister on the steps is a clip, a set of stills, and a transcript, and you want them together.

What the questions look like

  • “Every clip where this person is on screen, from anything we have shot.”
  • “The interview where the mayor talks about the flood defences, some time around 2015.”
  • “Photographs of the old market before it was demolished.”
  • “The rights letter for the archive footage we used in the anniversary piece.”

None of those has a keyword in it that a filename would contain. All of them are how a desk actually asks.

Finding it is half the job

This is where an agent earns the name. From any result you can move straight into asking DeepFile about it: what is said in this interview, who else appears, when was this taken. That is the difference between a search engine and something that behaves like a well-informed colleague, and it is usually what the person wanted in the first place — the answer, not the file.

What it will not do

DeepFile will not index your archive overnight; a large library takes time to read, and we run it as batch jobs rather than pretending otherwise. It cannot recover detail that a bad scan threw away. And it runs where your archive lives — on your own infrastructure if the material cannot leave the building, which for a publisher protecting sources is the normal case, or on ours if you would rather not run anything.

The archive is the asset. This is how it starts behaving like one.

The products this is about

Want the longer version?

Each piece is the short form of a conversation we have often. Tell us where you are and we will tell you what it would take.

Talk to us ↗