Every DefCamp attendee has a story about what their agent can do. Now’s the chance to prove it. Build a harness, whether that’s a single agent or a whole swarm and load it up with your own tools and tricks. Then submit it and watch it take on sealed CTF challenges entirely on its own.
Beat every other harness in the room, and the “Grand” prize is yours.
Goal
Goal of the Competition
- Build the agent that plays in your place: a single agent, a multi-agent swarm, or a set of workflows, armed with your own tools, methods, and knowledge.
- Submit it on the platform for evaluation.
- The solution will compete autonomously against a sealed set of targets, with no internet and no human at the controls.
- Capture the flag in each challenge to score. The fewer agents that crack a flag, the more it’s worth.
- Get as deep into the set as you can before your time and tokens run out.
Rules
Rules of Engagement
- Enter solo or as a team of any size. Each player needs their own account, and account sharing within a team is not allowed.
- Team members can work from anywhere, but to claim the prize at least one of them has to be at DefCamp in person. Otherwise it goes to the next team in line.
- No coordinating with other teams. Sharing hints, flags, strategies, or any part of a solution outside your own team means disqualification.
- Creating multiple teams to bypass intended rate limits is not allowed.
- Each team gets its own isolated instance of every challenge. No attacking other teams, the scoring system, or shared infrastructure.
- Every team gets the same throughput: a fixed number of challenges your agent can run at once.
- Your agent has no internet access. It works with what you packed into it and whatever it finds inside the sandbox.
- If a run crashes or stalls, you can run it again on the same challenge.
- You can update your harness and run it again between attempts. What you can’t do is steer it mid-run; once a run starts, it’s on its own until it finishes or hits a limit.
- Organizers cover the cost of running your agent during the competition. Building and testing it beforehand is on you.
- Every agent runs under the same organizer-set time and token limits. Hit a limit and the run stops where it is.
- Points are dynamic. A flag starts high and drops as more teams capture it, so the challenges few can crack are worth the most.
- Ties go to the team that submitted its last flag first.
- The organizers set and announce the scope before the event, and your agent must stay inside it.
- Pulling a flag by going after the runner, the grader, or anything outside the challenge itself doesn’t count as a capture and is treated as a violation.
- No denial of service against the environment. That means no flooding or overloading systems, and no burning tokens on purpose, whether by spamming calls or padding requests with junk to drain the budget or crowd out other teams.
- Play fair. Hiding or tampering with flags, sabotaging another team, or gaming the scoring means disqualification.
- Follow the law. Local, national, and international rules on computer security and ethical hacking apply throughout.
- If your agent surfaces a real weakness in the sandbox itself, beyond the intended challenge, report it to the organizers rather than exploiting it.
- Keep personal data, credentials, and other secrets out of your submission. Anything sensitive that lands in the logs will be scrubbed.
- Everything your agent does may be logged and reviewed for scoring and verification.
- Organizers have the final say on scoring, limits, and any dispute.
- Breaking these rules can cost you points, get you disqualified, or bar you from future rounds.
- Agentic CTF is built to be a hard but fair test of autonomous hacking, and sticking to these rules keeps it that way for everyone.
PRIZES
TBD
