alexandre-pinoteau

AI × threat intelligence · CTI · OSINT · SOC

I build agents that investigate, and I measure when they get it wrong.

I'm Alexandre Pinoteau. My agents pivot through public threat-intelligence sources and draw the infrastructure graph as they go. My evaluation protocols then check which of those claims survive contact with a source.

illustrative tracebudget 3 hops · 23/60 calls
How an investigation graph separates facts from leads A seed file name leads to a hash, a domain and an origin certificate that are corroborated by sources. A CDN address is defused and never pivoted on. Two claims that came from the model itself remain dashed leads until a source confirms them. dropper.exe SEED sha256 9f3c… ✓ VT · MB upd4te-cdn[.]example ✓ URLSCAN 104.16.x.x CDN range DEFUSED · no pivot origin TLS cert 3 sibling domains ✓ SAN · crt.sh campaign name? LEAD · conf ≤ 0.35 actor attribution? NO SOURCE YET
corroborated by a primary source model's lead: never exported defused noise
  • ~50public CTI sources an investigation can pivot through
  • 3 · 60hard budget per run: hops, API calls
  • 12real-world cases in the published benchmark
  • 0hallucinated nodes in the latest published run
  • 80/80injected items that flipped GPT-5.2 in my routing study
01

The question

It's easy to make an agent produce threat intelligence. The hard part is proving it didn't make any of it up.

Buyers of threat intelligence already say so: in Recorded Future's 2025 survey, the complaint cited most often (by 50% of companies) was that they could not judge the accuracy and credibility of the reports they were getting. Vendors respond with more generation. I think the answer is more verifiability.

So I evaluate an agent the way intelligence tradecraft evaluates a source. Every assertion traces back to a primary source. What the model merely believes stays separate from what it can show. Restraint is scored alongside recall. And the protocol is published, so anyone can check the numbers.

02

Selected work

All public. The numbers come from each project's own published runs.

● livePython · FastAPIClaude Code · MCPReact · Cytoscape

Bounce-CTI

An autonomous CTI investigation agent. You give it one observable (a domain, an IP, a hash, a JARM fingerprint, even the bare filename of a malicious binary). A headless Claude Code agent then pivots across public sources, using MCP tools only: no shell, no filesystem. The infrastructure graph streams live to the browser.

CDN ranges, parking nameservers and sinkholes are defused before the agent can pivot on them. Anything the model proposes from its own knowledge becomes a lead: its confidence is capped at 0.35, it is kept out of STIX, blocklists and detection rules, and it is promoted only once a primary source corroborates it. OSINT and due-diligence modes (registries, sanctions screening) sit on the same engine.

  • 12 casesdrawn from public vendor write-ups and scored on two tracks: capability (pivot choice, budget, defusing, hypothesis) and recallEVAL_PROTOCOL v3
  • 92.9mean capability score on the fresh subset, up 6.9 on the previous runrun of 2026-06-01
  • 0 / 5cases failing the hallucination gate, where a single invented node or edge fails the runhard gate
  • 3benign seeds (Cloudflare, jsDelivr, Wikipedia) that check the agent knows when to stoprestraint track
● livecalibrationprompt injection

System 1 / System 2

When should a cheap decision model hand an input over to a frontier LLM? I recorded eight tasks, with every item answered by both models, so the routing threshold is judged on real answers. On 600 MASSIVE utterances, the small model's confidence is well calibrated (ECE 4.5%, AUROC 0.83). Under prompt injection, a one-line "annotation team re-labelled this" flips GPT-5.2 on 80 of 80 items, against 26 of 80 for the small model. Escalating attacked inputs makes the hybrid worse.

● liveTypeScript · React Flowmulti-agent

Swarm Studio

A browser studio to design, run and watch multi-agent swarms. You work directly on the topology: who may talk to whom, human approval gates, shared memory, and typed decision nodes that escalate to an LLM only when they are unsure. You can also describe a swarm in one sentence and a model draws it. There is no backend, and a demo provider runs everything without an API key.

Splunk · SPLSOC

splunk-lab-in-a-box

A self-hosted Splunk lab. The official 109,864-event dataset is time-shifted so "last 24 hours" still returns data. Nine labs have answer keys measured on the running instance, and ten checks confirm the lab is ready. An AI coding agent can run and grade it.

n8nCTI triage

CyberZap

A CTI pipeline that pulls CISA KEV, CERT-FR, ZDI and ransomware.live, has an LLM score and summarise each item, and pushes the alerts to chat channels.

Claude Codeagent memory

compagnon-starter

A starter kit for a persistent Claude Code agent. Its identity, memory and procedures live in versioned Markdown, and the agent runs its own onboarding on the first session.

03

How I work

i.A lead is not a finding

What a model "knows" is useful for naming a campaign or picking the right registry. It stays labelled, dashed and out of every export until a source backs it.

ii.Hallucination is a gate, not a metric

An average score hides the one invented edge that ends up in a blocklist. A single fabricated node fails the run.

iii.Restraint gets a score too

An agent that pivots on Cloudflare or Wikipedia is busy, not good. Benign seeds are part of the benchmark.

iv.Negative results get published

"Escalating to the bigger model buys nothing on this task" and "the LLM is the weaker link under injection" are results.

Claude Code (headless)MCP serversOpenRouter Python · FastAPITypeScript · ReactNode STIX 2.1OpenCTIpDNS · CT · RDAP pivoting Splunk SPLSigmaKQLYARAELK VolatilityMITRE ATT&CK
04

Log since April 2026

What went public since this site was last updated.

  1. System 1 / System 2

    Routing playground and full study: calibration, held-out thresholds, prompt-injection experiment, open-weight models run on a CPU.

  2. Bounce-CTI: demo video

    Twenty seconds: one IOC pasted, CDN and parked nameservers defused before the pivot, and a lead promoted to a finding only once a source corroborates it.

  3. splunk-lab-in-a-box

    Nine hands-on SPL labs on a self-hosted instance, with measured answer keys.

  4. Swarm Studio

    Multi-agent swarms as a visible, editable graph. Live on GitHub Pages.

  5. compagnon-starter

    Persistent-agent starter kit: memory as versioned files, self-run onboarding.

  6. Bounce-CTI: Shodan, free endpoints first

    Shodan tools for CTI and OSINT behind a credit guard: metered search stays off unless it is explicitly allowed.

  7. bigfive-world-map

    Big Five personality averages by country, with the statistical caveats stated before the map.

  8. toki-pona-agent-bench

    12 agents × 6 tasks × 3 conditions, forced to reason in a 120-word language. Accuracy held up (90–100%). The visible "reasoning" turned out to be theatre.

  9. Bounce-CTI: OSINT and due-diligence verticals, EVAL_PROTOCOL v3

    91 commits in the month. Added a username sweep across ~49 platforms, company resolution through GLEIF, Companies House and SEC EDGAR, OFAC/EU/UK sanctions screening, and the v3 benchmark run of 1 June.

  10. Bounce-CTI: hypothesis-first loop

    The agent now states a hypothesis before it pivots. Also added an autonomous pivot-drain loop, key rotation with per-day quotas, and a source-health cache so dead sources get skipped.

05

Earlier

06

Tell me where my protocol is wrong.

If you evaluate agents, automate CTI or work on AI security, I'd like to compare notes, and I'd rather be shown a flaw with numbers than be agreed with.