Flagship case study — Live
Blackcipher
Search Engine
A search engine we built ourselves — crawler, index, ranking and query layer — because relying on someone else's API teaches you nothing about how search actually works.
Overview
Every search product we used felt like a black box — and the black box was expensive. We wanted a system where we controlled the index, understood the ranking, and could explain to a client exactly why a result appears where it does.
So we built one. It crawls, indexes, ranks and serves results through a documented API, and it's the reference architecture behind much of our client work.
The Problem
Off-the-shelf search APIs are fast to integrate and hard to live with. Three problems kept surfacing:
- Cost that scales against you. Every query is a bill, and the price grows with your success.
- No control over relevance. If the ranking doesn't fit your domain, you get a settings page, not a solution.
- No understanding underneath. When results are wrong, there's nothing to debug — only a vendor to email.
We wanted the opposite: full ownership of the pipeline, tunable relevance, and a cost curve we understood.
Our Approach
We treated search as four separate problems and made each one replaceable, so improving ranking never risked breaking crawling.
- Crawl: a queue-driven fetcher with politeness rules, retries and duplicate detection.
- Process: text extraction, language handling, tokenisation, stop-word removal and stemming.
- Index: an inverted index with positional data, plus fast lookup structures for prefixes and phrases.
- Rank: a scoring model combining term frequency, field weight, recency and link signals — each configurable.
Two rules ran through the whole build: every component exposes metrics, and every ranking decision has to be explainable to a human.
The Solution
The result is a working search platform with a clean API in front of it. A developer can index a corpus, query it, and get scored results with highlighted snippets back as JSON.
Because the query layer is the same for us and for clients, anything we learn on the engine transfers directly into client projects: search inside a store, a document management tool, a support knowledge base.
Key Features
- Relevance-ranked results with tunable weights
- Sub-200ms median response on the reference corpus
- Snippet highlighting and result previews
- Prefix and phrase matching with typo tolerance
- Filters and facets by source, type and date
- Documented REST API plus a lightweight web client
- Ingest pipeline for PDFs, HTML and plain text
- Usage metrics, query logs and slow-query reporting
Tech Stack
[EDIT] Adjust this list to match what you actually built.
Results
[EDIT] Replace every figure above with your measured numbers. If you can't measure it, don't claim it.
Screenshots
[EDIT] Swap in real screenshots of the engine (blur or crop anything confidential).