Once a week, an agent at StackOne goes hunting for new prompt injection attacks. It reads papers, trawls Reddit, tries what it finds against real models, keeps the attacks that land, and retrains StackOne's defence model on them. A person still approves every deployment. Guillaume Lebedel, their CTO, puts the whole thing on screen, including the five experiments out of six that failed.In this episode of In The Loop, Guillaume talks about how and why they have built an auto-research loop. What one is in plain words, how you pick a goal ai agents can measure, how you sample 100,000 test cases down to something you can afford, and why he runs evals on cheap models before trusting anything. We also spoke about token leaderboards and the impact its had internally. ⏭️ Episode highlights(05:41) – What an auto-research loop is(11:53) – Why an agent is a folder(14:47) – Screen-share: the attack-hunting agent(18:00) – Six experiments, one promoted(27:29) – Sampling 100,000 test cases down(31:48) – Ninety per cent of tokens are wasted(39:33) – Turning off extra usage the same day(41:13) – Screen-share: the token derby
We do not know your name or where you live, but the map says you were here. Thank you for stopping by — every flag below is someone who came to say hello.