Back to all updates

about 2 months ago

Announcing: Find Evil! Hackathon Finalists

THRILLED to announce the five finalists for the Find Evil! AI Hackathon!

The mission: build an autonomous AI agent that investigates a compromised system the way a senior incident responder would, and prove every finding it makes. 4,413 registered. 291 teams submitted a working agent.

Winners announced on a live webcast Wednesday, August 19, 12:30pm EDT! Put on your calendar now: https://www.linkedin.com/events/7493520348248969216/

The five finalists:

Camel built by Allister Beharry. Camel lets the AI write and run its own forensic code inside a sandbox, and every finding traces back to the exact commands that produced it. Judge: "Could see the autonomous in action. Constraint implementation and audit trail are present and visible in the report itself. All checks passed." Devpost: findevil.devpost.com/submissions/1050173-camel or GitHub: github.com/allisterb/Camel

FindEvil built by Marshall Lee: An autonomous Linux incident response agent with write protection enforced in the code itself, and it posted 98.6 percent recall across a 552-attack test harness. The team also found and fixed two security flaws in its own guardrails during development, and published the fixes. Judge: "This is really good and one of the few approaches able to deal with Linux evidence." Devpost: findevil.devpost.com/submissions/1042619-findevil-autonomous-linux-ir-forensics-agent or GitHub: github.com/marlyocat/findevil

Mulder built by Caleb Evans. It runs a five-phase investigation across memory, disk, network, mobile, and logs, and in one run it logged 773 tool calls across 11 systems and 120 GB of data. An adversarial phase works to disprove its conclusions before you ever see them. Judge: "The hallucination precautions and sources being required and challenged are fantastic. He challenged the tool to work harder before simply jumping to conclusions." Devpost: findevil.devpost.com/submissions/994654-mulder or GitHub: github.com/calebevans/mulder

Protocol SIFT++ built by Zheng. It is read-only by construction with a tamper-evident chain of custody, and a component called the Skeptic reruns its tools to disprove its findings. During the validation round, it retracted a rootkit finding it had previously confirmed. Judge: "Very good entry. Excellent usability, was able to get up and running quickly." Devpost: findevil.devpost.com/submissions/1049723-protocol-sift or GitHub: github.com/tupils1/protocol-siftpp

TRUDI built by Trinity Harrison. It pairs an autonomous investigator with a second AI that challenges what the first one concludes. When it was handed a briefing containing planted indicators, TRUDI ran the tools, decided the briefing was wrong, and refuted its own starting assumptions. Judge: "Self-correction at call 54 is genuine. It ran a command on the briefed indicators, got empty, and refuted its own briefing. Devpost: findevil.devpost.com/submissions/1019352-trudi-threat-response-unit-for-digital-investigation or GitHub: github.com/nebulae/trudi

__

The judging: 90 working incident responders ran 1,775 evaluations. They ran the agents against case data, and then they attacked them with path traversal, command injection, and attempts to make the agent alter the very system it was investigating. Every finalist was tested again against evidence the teams hadn't seen. More about their validation process (a heroic effort on it's own) on the live webcast.

CONGRATS to all five teams. To every judge who gave hours and effort, to every one of the 291 teams who submitted, and to SANS Institute for sponsoring, THANK YOU!

Rob T. Lee 
SANS Institute