I'm releasing today TarantuBench-v2, a collection of over ten thousand web-app ctfs. They are synthetically generated, verifiably exploitable, and include a two-tier detection mechanism that attempts to flag when an agent finds an unintended solution.
It is a follow-up to v1, which included one hundred, and which were mostly useful for benchmarking.
With ten thousand labs, you can:
- Evaluate different deployments of different harnesses you might be using
- Compare and contrast different underlying models
- Train existing models and agents
This effort is a work-in-progress, in which I'm trying to synthetically generate increasingly sophisticated CTFs, in high volume, and with improving detection capabilities of shortcuts that an AI might find.
The dataset and all the technical explanations are available on huggingface.
[link] [comments]
from hacking: security in practice https://ift.tt/YP1XsW2
Comments
Post a Comment