From .onion to Structured Data: A Hands-On RansomLook Workshop — Olivier Ferrand

Join us at hack.lu 2026 — Info & Registration

Duration: 90 min

Type: Workshop

Speakers: Olivier Ferrand

Abstract

Ransomware leak sites are a noisy, hostile and short-lived source of intelligence: onion services that go down without warning, layouts that change overnight, and increasingly, anti-bot challenges designed to keep collectors out. RansomLook is an open-source platform that has been tracking these sites for 4 years, turning scattered extortion pages into structured, searchable and pivotable data used by CERTs, researchers and threat intel teams.

This hands-on workshop takes you from the outside in. We start with a tour of the platform and the kind of pivoting it enables, then move to the collection layer: how a scraper is configured, what breaks in practice, how to write an init script that gets you past a CAPTCHA-protected leak site, and how to write a parser that turns a raw page into victim records the platform can ingest.

Description

What this workshop is:

RansomLook (https://www.ransomlook.io) monitors ransomware and data-extortion leak sites, normalises what it finds, and exposes it through a web interface, an API and notification feeds. It currently tracks ~600 groups across ~3400 sites. This session is not a product demo: it is a working session on the part of the project that is hardest to get right, the collection layer, and it is built so that participants can contribute to it afterwards.

Who it is for:

CERT and CSIRT analysts, threat intelligence practitioners, and researchers who collect from onion services or who want to. Anyone who has written a scraper that worked for three weeks and then silently returned empty pages will recognise the problems we work on.

Outline (90 minutes):

RansomLook: origins, design decisions, and the team : why the project exists, what it deliberately does not do
Platform tour and pivoting : walking real cases: from a single victim mention to group infrastructure, timelines and overlaps
Environment check : everyone running locally before we go further
Scraping: configuring a collector properly : site definitions, Tor plumbing, scheduling, retries, and reasons a scraper silently stops working
Exercise 1: writing an init script to get past a CAPTCHA : session bootstrapping against a challenge-protected leak site, on a captured replay so the exercise does not depend on the site being up
Exercise 2: writing a parser : turning a raw leak-site page into structured victim records: selectors, deduplication, and what to do with the mess that does not fit the schema
Contributing back : the parsers currently missing, and how to open a pull request

Prerequisites:

A laptop able to use VM (we work with qcow2 disk) . Basic Python is enough, participants write short scripts, not framework internals. A pre-built environment and an offline corpus of captured leak-site pages are provided, so no exercise depends on conference Wi-Fi or on a live onion service being reachable. A shared hosted instance is available as a fallback. Setup instructions are published fews before the conference on the RansomLook Projet repository.

What participants take away:

A running local RansomLook instance, an init script and a parser they wrote themselves, an understanding of how a leak-site collection pipeline fails in production, and a concrete way to contribute to a project their own team may already be consuming data from.

A note on scope and ethics

The workshop deals exclusively with fake DLS.

View on pretalx