Home / Work / Case study

Flagship case study — Live

Blackcipher
Search Engine

A search engine we built ourselves — crawler, index, ranking and query layer — because relying on someone else's API teaches you nothing about how search actually works.

Overview

Every search product we used felt like a black box — and the black box was expensive. We wanted a system where we controlled the index, understood the ranking, and could explain to a client exactly why a result appears where it does.

So we built one. It crawls, indexes, ranks and serves results through a documented API, and it's the reference architecture behind much of our client work.

100k+Pages indexed in the reference corpus
<180msMedian query response time
04Ranking signals, fully tunable
99.9%Uptime across the last quarter

The Problem

Off-the-shelf search APIs are fast to integrate and hard to live with. Three problems kept surfacing:

  • Cost that scales against you. Every query is a bill, and the price grows with your success.
  • No control over relevance. If the ranking doesn't fit your domain, you get a settings page, not a solution.
  • No understanding underneath. When results are wrong, there's nothing to debug — only a vendor to email.

We wanted the opposite: full ownership of the pipeline, tunable relevance, and a cost curve we understood.

Our Approach

We treated search as four separate problems and made each one replaceable, so improving ranking never risked breaking crawling.

  • Crawl: a queue-driven fetcher with politeness rules, retries and duplicate detection.
  • Process: text extraction, language handling, tokenisation, stop-word removal and stemming.
  • Index: an inverted index with positional data, plus fast lookup structures for prefixes and phrases.
  • Rank: a scoring model combining term frequency, field weight, recency and link signals — each configurable.

Two rules ran through the whole build: every component exposes metrics, and every ranking decision has to be explainable to a human.

The Solution

The result is a working search platform with a clean API in front of it. A developer can index a corpus, query it, and get scored results with highlighted snippets back as JSON.

Because the query layer is the same for us and for clients, anything we learn on the engine transfers directly into client projects: search inside a store, a document management tool, a support knowledge base.

Key Features

  • Relevance-ranked results with tunable weights
  • Sub-200ms median response on the reference corpus
  • Snippet highlighting and result previews
  • Prefix and phrase matching with typo tolerance
  • Filters and facets by source, type and date
  • Documented REST API plus a lightweight web client
  • Ingest pipeline for PDFs, HTML and plain text
  • Usage metrics, query logs and slow-query reporting

Tech Stack

PythonNode.jsExpress MongoDBRedis queueREST API DockerNginxVanilla JS client

[EDIT] Adjust this list to match what you actually built.

Results

100k+Documents indexed and searchable
<180msMedian query latency
~65%Lower search cost versus the third-party API it replaced
04Ranking signals tuned for relevance, up from zero control

[EDIT] Replace every figure above with your measured numbers. If you can't measure it, don't claim it.

Screenshots

Search results page with a search bar at the top and a list of ranked result rows
01 — query view · ranked results with snippets
Dashboard view showing indexing metrics and a table of crawled sources
02 — index health · crawl and coverage metrics [EDIT]

[EDIT] Swap in real screenshots of the engine (blur or crop anything confidential).

Build something like this

Have a system
that needs to exist?

Search, dashboards, internal platforms — if someone else's tool doesn't fit, we'll build the one that does.