Timnit Gebru (she/her).

@timnitGebru@dair-community.social

Perfil original

I'm posting here what I had written to a news outlet in April during the whole Mythos hype 🙄 and people keep on being like BUT why are you not making CLEAR arguments against Bla BLa Bla? Cause I'm learning that its a trap to make me waste my time debunking claims while they get trillions to implement their dystopian future and we're stuck talking about whatever they're doing rather than implementing our visions. 🧵

My concerns were on the types of claims made about Mythos. There has been a growing trend of companies making claims and these claims being repeated without any way to verify them or having misleading ways of evaluating their claims. Each time we have had the actual code base, data and models to analyze, we have found a lot of issues with claims in many different domains of machine learning. My own work has shown this in other scenarios.

For Mythos, for example, Anthropic doesn't tell us the number of false positives their tool returns, i.e. the number of times their tool says that something is a vulnerability and it ends up not being. My security expert collaborators tell me that this is one of the most important metrics by which security tools are judged, because it tells you the difference between a useful tool and a useless one that engineers won't use.

Anthropic also claims that Mythos can replace security experts. It's one thing to claim that you've built a useful tool, another to claim that you can replace experts. Some security experts have even said that it's dishonest to say your tool is superior to security experts because it found bugs in old codebases and we don't know how often people audit them for bugs and fix them. Again a misleading claim that is repeated by those outside of the company.

Another thing that I find funny is that Anthropic is making all these claims about security while they themselves couldn't stop their own source code from leaking. And from analyzing that code, people have seen how things are done now with "vibe coding" using Claude. Even things that can be done simply are done with brute force, i.e. trying all possible scenarios before arriving at an answer or solution.

How much computational power does it take to find any of these vulnerabilities Anthropic says they found? How much money would someone have to pay to use their tool vs bonafide experts?

A lot of security and safety is about processes, checks, clear people in charge of clear things, and clear access limits. That goes out the window with these “agents. It becomes harder to identify where things went wrong, or what types of tests you need to do to check for vulnerabilities and who is in charge of which issue.

Some experienced software engineers talked about how it's easy to just press “accept” of the buggy code that you see generated by Claude. So the question is, what about the new vulnerabilities created by using these “agents”? We’ve seen so many examples of issues, like people wiping out their entire production data.

Impulsos
Me recuerda a la urgencia por sacar producto: en una cooperativa pequeña de software libre, la tentación de aceptar lo que genera el asistente es fuerte, pero el coste lo pagas luego. Nosotras hemos empezado a tratar esas sugerencias como código de terceros: revisión, pruebas y documentación antes de tocar producción. Así no renuncias a la herramienta, pero la integras con criterio comunitario.
16 de junio de 2026

Scientific American's Megha Satyanarayana asked me what advice I have for young scientists. My answer:

"Question everything, question the source of information that you get [...] Just follow all the rabbit holes [...] You spend years and years and years still asking the same question or following a specific rabbit hole that other people are too tired to follow. And that’s really where there is no shortcut to scientific discovery or innovation. You have to struggle."

scientificamerican.com/article