FLAWED's Flaws and What This Means for Industry Research
FLAWED’s Flaws and What This Means for Industry Research Disclaimer: The views expressed here are my own and do not represent those of any current or former employer or affiliated organization. On September 17th, I quote tweeted Trail of Bits’s blog post titled “1Password's AI patching benchmark is misleading,” which also referenced Davi Ottenheimer’s “Disinformation Pushed by 1Password: Their AI Patching Report is False.” Both criticized “Frontier Models’ Vulnerability Patches are Often F.L.A.W.E.D” (henceforth referred to as “FLAWED”) from 1Password's Off‑by‑1 Labs. I saw FLAWED when it was released and discussed it with other researchers; we classified it as slop and moved on. What I had not realized at the time was how far 1Password’s distribution had carried it: into news coverage and defender roadmaps. Watching this work obscure more rigorous research from less-resourced groups compelled me to post on Twitter, and the responses to that compelled me to write this blog post. The original thread What follows is the thread, including two follow-up replies, reproduced verbatim. There are glaring issues in this paper beyond the evaluation problems raised by Davi and ToB, including ones that make me question the ratio of human to AI assistance here, but a few are so egregious we should consider what research norms we are demanding from industry labs. While these errors include incorrect diagrams (e.g. where is the asterisk figure 1 claims are on the relevant steps?), arithmetic errors (e.g. 2.8 != roughly 4), and internal textual inconsistencies that can be gleaned from a skim, the citation problem is the most serious. Industry papers may reasonably have somewhat lower citation density. But FLAWED adopts the form and rhetoric of rigorous research (even claiming there is “very little prior work” in this specific area whatsoever), citing only 19 sources (mostly corporate blog posts plus XKCD).! That claim does not match the literature. The concurrent PatchBench paper has 73 citations (https://arxiv.org/pdf/2609.04075) mainly of academic research papers. As an example, Meta's AutoPatchBench is not cited, despite being obvious prior work from prominent researchers. FLAWED did not cite and did not read this NDSS paper (https://ndss-symposium.org/ndss-paper/chasing-shadows-pitfalls-in-llm-security-research/) about pitfalls in LLM security research, including the relevant subfield, which would have really benefited the work because FLAWED has such identified pitfalls. There is also a factual issue around Patch the Planet. Calif and HackerOne participated in “vulnerability triage, coordinated disclosure, and additional focused vulnerability discovery efforts,” according to OpenAI, yet the paper only points to ToB engineers. I understand why ToB gets the attention here; they were OpenAI’s partner and all of the authors’ most recent former employer. Being former ToB myself, I know it's easy to cite what's top-of-mind. But the statement remains inaccurate and inappropriately so for a research paper. I’ve done industry research that became blog posts and whitepapers, and I’ve published academic work. The conventions of the format create corresponding expectations. This claims and adopts the tone and format of authoritative research without adopting the necessary rigor. Mistakes are normal; that’s part of why we publish, reproduce, and critique research. But it’s also why corrections, and when needed: retractions, exist. I hope that 1Password issues a retraction or a correction. I hope they, minimally, partner with academic researchers, hire folks with research backgrounds, or engage with the community to prevent releasing work with such a high concentration of consequential errors again. I say this as a 1Password user who wants them doing this work. AI security needs research from groups independent from the frontier labs, especially work scrutinizing potential marketing claims. However, that makes enforcing strong research norms even more important, not less. While 1Password’s Off-by-1 Labs is not the worst offender in the world, I had hoped they would produce work of quality and integrity given that, if anything, their interests should favor credible research over slop. In response to someone bringing up frontier labs, I wrote: Of course! If you scroll through the rest of my tweets and retweets, you can see that both I individually and research I've boosted has been plenty critical of the frontier labs. But posting slop under company branding is just another way to not hold tech companies accountable. The fix is honesty and rigor, not plagiarism by omission and slop. We should amplify real, good, independent work from places like academia or EleutherAI, who have far fewer resources and less reach than industry labs. Omitting directly relevant academic work deepens that imbalance. 1Password's work gets the attention, press releases, etc. and the real thinkers don't. When some errors are obvious on a skim, the lack of care is extremely hard to excuse. 1Password is not some poor underdog!! I do want OpenAI's claims really evaluated! OSS deserves resources directed toward what actually makes it safer. CISPA, UMD, and Drexel did work on similar questions that is rigorous and also critical of LLMs for security that 1Password perhaps should have helped fund instead. In response to concerns about the quality of my rebuttal, I also wrote: I’ll call out malpractice and slop. I’m very curious what the truth is and look forward to real, rigorous work. Feel free to comment on the evaluation rebuttals linked. But proving a claim false does not require also finding the correct one; that asymmetry is how research works Afterward, CISOs, academics, industry researchers, and practitioners alike messaged me to say they had noticed the same problems. Many had not raised them because they lacked a forum or feared backlash. Those responses also exposed second-order costs, especially for academics, that may be less obvious to folks unfamiliar with the social systems surrounding research. Research should not be treated like sports 1Password employs many talented individuals and has a reputation for high quality work. Despite this, FLAWED has serious issues that must be raised. If you believe these issues should not be raised because FLAWED was critical of OpenAI, please know that research is not sports. OpenAI is not the Spurs, 1Password is not the Knicks, and Off-by-1 Labs is not Jalen Brunson. All research should face healthy skepticism. Rigorous evaluation of frontier-lab claims is both possible and necessary, including the claims inherent to Patch the Planet. EleutherAI regularly publishes precise work in this realm. The AI Now Institute delivered a sound write-up critical of Patch the Planet (even though it's not a research paper, it has more citations than FLAWED). Research should not become an influence operation FLAWED confirms a preconceived notion many already held, which likely explains some of its traction. But we cannot believe and amplify research merely because it matches our priors. The purpose of rigor is to protect us from conclusions we want to accept but that are not true. What is the worst-case scenario if we do not enforce this norm? This opens the door for malicious actors that could repeatedly select conclusions that flatter their institutional or personal interests, apply weak methods, and use corporate distribution and social engineering to make those conclusions disproportionately influential. I am not alleging that this happened here in any way, shape, or form. I am describing a vulnerability in our research ecosystem that exists independently of this particular paper or its authors’ intentions. These discussions are integral to fostering high integrity research environments. As Carlini writes in “Why I Attack,” “you can't fix something if you don't know it's broken, and so someone needs to show what's broken.” A paper is not a blog post Security research encompasses a broad range of activities: vulnerability discovery, exploit and tool development, threat intelligence, and empirical study. Empirical, scientific claims about human-AI interactions require study design and evidence beyond what is needed to demonstrate a vulnerability. The same distinction applies to genre. A blog post can present data-driven observations without claiming scientific authority. A paper presented as scientifically rigorous and peer-reviewed assumes additional obligations. FLAWED actually uses the phrase “peer review” to describe review by three industry peers thanked by the author, not the scholarly peer-review process. While citation count is not itself a measure of rigor, it is a useful proxy for whether authors accurately assess their claims and responsibly engage with the relevant literature (it’s early in the standard method for reading research papers for a reason). Plagiarism by omission is serious. I would be suspicious of any research that does not cite foundational papers in the area (e.g. this paper is considered foundational to AI patching research) as it indicates either a deep unfamiliarity with the literature or the aforementioned plagiarism by omission. I truly believe this work could have made a compelling blog post if the focus was on the actual technical contributions. “Codex (and GPT-4) can’t beat humans on smart contract audits” illustrates the path that could have been taken. This blog post makes an empirical claim about human-AI interactions, but explicitly states “Our assessment does not meet the rigors of scientific research and should not be taken as such. We attempted to be empirical and data-driven in our evaluation, but our goal was … not scientific publication.” A press cycle is not a publication process The form factor of both the release and update is misaligned with how research should work. Research needs a public record. Quoting Hillel Wayne on science, “One thing lay folk don’t realize is that science is social. For all we focus on “objectivity” and “evidence”, it takes place in human institutions and relies on how humans work. We care about integrity, trustworthiness, and reputation. While this often surprises outsiders, it’s ultimately necessary for science to work in the large.” Terence Tao makes a related point about math (critiquing OpenAI): “I have grown weary of this new practice of using press releases or social-media posts to communicate mathematical results”. As far as I can determine, FLAWED was not posted to a scholarly repository or public review forum. While doing so would not have constituted peer review, it would have created a conventional, versioned record. Industry research can be released as non-scholarly white papers, but FLAWED adopted the veneer of rigorous, scholarly research and was promoted publicly and privately as such. As far as I can determine, FLAWED was not posted to a scholarly repository or public review forum. While doing so would not have constituted peer review, it would have created a conventional, versioned record. Industry research can be released as non-scholarly white papers, but FLAWED adopted the veneer of rigorous, scholarly research and was promoted publicly and privately as such. The update does not address the rebuttals. It is verifiably untrue that the only issue was that “the subset of variables used in our approach over-constrained the output.” If that genuinely was the only problem with this, if the work was merely incorrect due to evaluation errors instead of being reasonably classified as slop by many qualified individuals, my blog post would read much closer to this critique. A public claim should have a public correction. Given the substance of the rebuttals, it is inappropriate for the only feedback channel to be a private email, and it is dishonest to keep the preprint as it stands without a retraction. This information must also be shared with the journalists that promoted this research. Carlini’s response to InstaHide models what researchers should do when broken work continues to be promoted. Integrity matters in research Why do I care so much? Because dishonest research does more than put a false claim into the literature. It distorts research agendas. Work that seems duplicative or incorrect according to prior research is often not funded. Researchers abandon problems that appear to have been solved or pivot after being falsely scooped. It wastes research labor. Researchers spend time reproducing or extending misleading findings, reconciling sound results with them, and reviewing work built on false premises. It cannibalizes attention and pollutes the epistemic environment. Misleading claims crowd out more rigorous work, contradictory evidence is discounted, correct conclusions are delayed, and corporate reach becomes a proxy for validity. Unfortunately, FLAWED has already led to this (that was the subject of multiple messages I received after posting on Twitter; a reason for this very post is to help academics currently in the middle of these conversations). We need to foster integrity in our community. Luckily, although this work reached people through news articles and social media, it does not appear to have directly entered the academic literature yet (some seasoned researchers told me they spotted the problems immediately). That is reassuring, but it is not foolproof. The onus is also on all of us to read the work we cite and amplify rather than treating its format, affiliation, or existing citations as proxies for validity. We should care about OSS security I strongly believe that any company with a strong security program should have a robust security research function, and I also believe that I have a responsibility to shape and enforce the norms of my community, especially with respect to surrounding power dynamics. I am genuinely seeking evidence surrounding interventions that improve open-source software security. Given a fixed budget, which interventions work best: hiring more maintainers, providing existing maintainers with token spending, partnering with consultancies who are not allowed to use AI, or something else entirely? This question is still open to the best of my knowledge; FLAWED diverted time and resources away from efforts that might genuinely answer it. If you are conducting real, rigorous research in this area that is not receiving sufficient attention, please contact me. I will review it and if I feel that it is strong, I will share it widely and do what I can to help it reach the attention and resources it deserves.