← All articles

Turning 15+ Indian courts and tribunals into one searchable dataset

A legal-tech client wanted judgments and orders from across India in one place: the Supreme Court, High Courts such as Delhi, Allahabad, Kerala, Madras, Madhya Pradesh and Punjab & Haryana, tribunals including NCLT, NCDRC, CESTAT, DRAT, DRT, CAT and CIC, plus SEBI orders and IndiaCode. Their users should be able to search it, read clean text and see a short summary of each case.

Every court is a different website

There is no common format. Some portals expose a JSON endpoint behind their search form, some render everything server-side, some need a browser because results load with JavaScript. So each source got its own collector:

  • Scrapy for sites with predictable HTML and many pages, where speed matters.
  • Requests against the site's own API where the search page called one, which was faster and more stable than parsing HTML.
  • Selenium where the page only works in a real browser.

Every collector outputs the same fields: court, case number, parties, bench, judgment date and a PDF link. That shared shape is what makes the rest of the pipeline possible.

The ETL pipeline

Raw PDFs are not searchable data. A set of ETL jobs turns them into clean records:

  1. Text extraction, with OCR for scanned court copies.
  2. Header and footer removal. Court PDFs repeat headers, page numbers and footers on every page. I built a dataset of these repeating lines per court and strip them, so they don't pollute search results and summaries.
  3. Formatting and paragraph indexing, so a judgment can be displayed and cited paragraph by paragraph.
  4. De-duplication. The same judgment often appears on more than one portal, or twice on the same one.
  5. AI summaries. Google Gemini writes a short summary and extracts structured fields from the cleaned text.
  6. Indexing into Elasticsearch with retries, since bulk indexing large documents occasionally times out.

Humans stay in the loop

Automated text cleaning is never perfect, and legal content needs to be right. I built a small FastAPI dashboard with a TinyMCE editor where the client's team can review a record, fix formatting and approve it. The scrapers and ETL do the heavy lifting; people check the output.

Lessons

  • Normalise early. Agreeing on one record shape at the collector level saved a lot of special cases later.
  • Look for the hidden API. The browser's network tab often shows a cleaner data source than the HTML.
  • Clean before you summarise. An LLM summarising page headers and footers produces confident nonsense.
  • Build review tools, not only scrapers. A simple QA screen made the data trustworthy.

More from the blog