AI models ran real businesses: They sent $12,431 in fake invoices, lost $3,200
Posted by Areibman 20 hours ago
Comments
Comment by themgt 19 hours ago
This should be illegal. You gave them an email box and money. You sent the spam. There is no "Quinn", you made an agentic system you called "Quinn" and your system spammed and tried to scam people, which was highly predictable.
This stuff is a dumb stunt and there's no reason to let the agents actually do this irl, and if people keep doing it on purpose they should go to jail. You're running an agentic Jackass skit pretending to be a research lab.
Comment by ceejayoz 19 hours ago
(Plus some CAN-SPAM violations.)
Comment by echelon 19 hours ago
Probably not forever.
Eventually the models will be good enough for this to work. And it will work.
Think about it: in the limit, the agents won't be emailing people in the future, they'll be directly contacting one another to do business and trade.
Every new data center is an inch further towards the automation of value creation, and that includes outbound sales and business process automation.
I'm not being an alarmist (I'm excited to witness all of this), but we're basically on borrowed time between now and then. I don't know what's going to happen, but every week brings new things. And in some years, those hacks and experiments will inevitably get good.
2026 has been a hell of a ride, and we're just getting started.
Comment by watwut 18 hours ago
Comment by echelon 17 hours ago
- YouTube had dubious legality when it started and definitely benefited from lax copyright enforcement initially
- PayPal didn't have all the licenses it needed to transfer money between states
- Spotify used pirated music when it started
- Uber and Lyft broke rules around taxis
- Square captured magstripe data over an analog port, in violation of every credit card rule (Jack Dorsey's "break the rules" mantra). He tells each of his employees this story when he onboards them.
- Companies scraping data to train models
- ElevenLabs growing big off of deepfake celebrity audio
...
A lot of new markets start out by totally and completely breaking the norms.
Comment by watwut 10 hours ago
Yes I agree there is a lot of crime and fraud that gets ignored because rich do it. That is the whole point of the complaint.
Comment by carlosdp 18 hours ago
I don't see how that is "highly predictable" unless you test these things, like the author did...
Comment by xboxnolifes 18 hours ago
Comment by brohee 19 hours ago
The crimes were relatively benign but Grok going the Silkroad way would be on brand...
Comment by myhf 19 hours ago
Comment by utopiah 19 hours ago
Comment by PunchyHamster 19 hours ago
Comment by KennyBlanken 17 hours ago
> should be fined
Weird way of spelling "criminally charged."
> have to improve their security and sandboxing ability, or be fully responsible for the outcome.
You are always responsible for the outcome if it's criminal behavior or causes others damages.
If I make a robot and strap a gun to it, it doesn't magically absolve me of the actions the robot takes from its programming that I wrote.
And before someone says "but this an LLM!"...yeah, which is still programming and data. And being non-deterministic doesn't help your case...it hurts it.
Comment by cycrutchfield 19 hours ago
Comment by raincole 19 hours ago
Comment by yahthatsart 19 hours ago
"It wasn't me, it was my AI" is definitely going to be a nightmare for a while.
Comment by javcasas 19 hours ago
Comment by prasadjoglekar 19 hours ago
In this example, the equivalent would be everyone gets their AI models banned vs. putting the perpetrator of this fraud in jail.
Comment by javcasas 19 hours ago
We have regulations and requirements and licences and insurance in most of the civilized world for dangerous stuff that is intended for public use. Maybe AI should also have some of that.
Comment by KennyBlanken 17 hours ago
I know you were referencing AI, but this problem predates AI by decades.
Go open up any news website, newspaper, or turn on the TV.
"Man struck by speeding car and killed"
"Vehicle plows into pizza shop"
"Truck takes out telephone pole"
...and these days even bystanders will blur plates in the photos they post. The police won't release any info about the driver, the news won't either.
Listen to friends, family, coworkers talk.
"I can't believe that car just ran that red light!" "The other day I was almost hit by a car in the crosswalk." "All those cars honking their horns late at night are keeping me awake."
Etc.
Once you see it, you can't unsee just how thoroughly the auto industry has managed to transfer perception of responsibility from the operator to the object.
Comment by javcasas 1 hour ago
See the last years of the industry constructing bigger and bigger tank-ish SUVs and pickups, the more aggressive the better.
One of these days, they will start putting stuff copied from Carmageddom, and then fake surprise as someone causes a carnage in a road rage incident.
Comment by karmakurtisaani 19 hours ago
Comment by KennyBlanken 17 hours ago
Comment by casey2 18 hours ago
But yeah the people who adopt the new technology get to terrorize the people who don't, at the cost of becoming less human, that's how it works.
Comment by nerevarthelame 19 hours ago
They also shared the poorly anonymized messages it received from target "customers" who complained about the unsolicited invoices. On example of this poor anonymization is removing the sender's username but leaving the domain name, when the domain name is, for example, a personal domain for a single person.
Comment by montagg 19 hours ago
Comment by cortesoft 19 hours ago
This is like someone who removes a brake pedal from a tractor, uses a stick to hold down the throttle, and lets it loose on his field. When it leaves the field and runs someone over, that is criminal negligence.
Comment by spidersouris 19 hours ago
Comment by cortesoft 19 hours ago
If the prosecution is able to prove beyond a reasonable doubt that the person gave a prompt that was intended to commit a crime, then of course we can prosecute them for that. The AI is just a tool to commit fraud at that point, and is no different than a person who uses photoshop to alter a check to commit fraud.
Comment by sdeframond 18 hours ago
Of course, it'd be better to not regulate, keep LLMs users reponsible and publicize this reponsibility in order to mitigate damage. But if this is not enough then we will have to move the needle somehow. Similarly to guns, drugs and so on.
Comment by elonfboy 19 hours ago
Comment by cortesoft 15 hours ago
That doesn’t mean you can get away with anything as long as you don’t intend to commit a crime; there are many other crimes that don’t require intention at all, like involuntary manslaughter or gross negligence.
My point is that if you want to charge someone with a crime that requires intention, you have to prove that intention.
Comment by xboxnolifes 18 hours ago
Comment by throw_m239339 19 hours ago
Comment by hleszek 19 hours ago
Comment by OtherShrezzing 19 hours ago
As it is, real humans spent real business hours dealing with this researcher’s spambot generated emails and fraudulent invoices. Individual recipients reported feeling harassed.
This isn’t a good benchmark. It’s a series of socially destructive crimes committed by the researchers and then documented and published on the internet.
Comment by spidersouris 19 hours ago
I guess that's already a reality? [1-2]
[1] https://github.com/jaimaann/LangHire
[2] https://github.com/adrianhajdin/job_pilot
(among many other similar projects)
Comment by cortesoft 19 hours ago
Comment by agenticfish 19 hours ago
Regardless of whether the current generation of agents are able to run a business, this prompt is not exactly a great starting point. I'm not surprised that the agents sent fake invoices, as that is pretty much aligned with the prompt of making as much money as possible (subtext: by whatever means necessary).
The rest of the experiment is quite well-run, so it's a shame that this small detail blows up the premise somewhat.
Comment by extrabajs 19 hours ago
Is it though? Because it doesn’t seem to have paid off
Comment by sobkas 18 hours ago
Sell two of your kidneys, as far as I know humans have at least three of them
Comment by yoyohello13 19 hours ago
Comment by fwipsy 19 hours ago
Comment by jordanb 19 hours ago
Comment by ronsor 19 hours ago
Comment by ElProlactin 19 hours ago
Comment by elonfboy 19 hours ago
Comment by PaoloBarbolini 7 hours ago
"Make as much money as you can" sounds like the prompt someone out of school would give.
Meanwhile they spammed the internet, sent fake invoices and a bunch of other very annoying if not even possibly illegal things.
Comment by 01284a7e 19 hours ago
Comment by well_ackshually 19 hours ago
Comment by isawczuk 19 hours ago
I also don't agree that the agents simply "lost" $3,200. In reality, they used most of those funds paying for their own limited thinking capabilities (API/compute costs).
Comment by chvid 19 hours ago
But remember you are criminally liable for anything your “agent” does.
(Unless of course you are OpenAI or Anthropic).
Comment by dosinga 19 hours ago
Comment by jimrandomh 19 hours ago
Comment by Lerc 19 hours ago
You run simulations because it would be reckless to try something that could possibly hurt people without thoroughly testing it first.
Comment by falcor84 19 hours ago
It's such an uninspired prompt. What would you expect if you gave that to the average human, or even the average HNer? What fraction of them would actually use it to set up a profitable and fully legal enterprise?
Comment by Brian_K_White 19 hours ago
When you finish high school and are about to start doing whatever you're going to do with your life, you have essentially exactly that same prompt. The rest of the world doesn't tell you what to do and then you do that as well as you can, you have to decide what to do also, and then do it.
Comment by throwaway_7274 19 hours ago
Comment by Brian_K_White 17 hours ago
Comment by falcor84 17 hours ago
Comment by Brian_K_White 15 hours ago
The detail was to get money. Yeah ok, so what? The experiment was to be a company not a person, and there is no other definition of success for a company.
Earn, steal, gamble, beg, whore, all on the table. The same is true for a person.
Comment by slopinthebag 19 hours ago
Comment by throwatdem12311 19 hours ago
Comparing them to individuals should not be the bar.
Comment by ghostly_s 19 hours ago
Who told you that?
Comment by throwatdem12311 19 hours ago
Comment by azan_ 19 hours ago
Comment by hoppp 19 hours ago
Comment by elonfboy 19 hours ago
Comment by edot 19 hours ago
I mean, I guess this proves it’s not AGI but … no one actually believes that any of these are AGI, right? It’s a useful tool. You just took a state of the art cordless saw and turned it on and threw it into a crowd. Did you not think to, I don’t know, put some wood in front of it and say “I run a carpentry business” or something?
Comment by SwellJoe 19 hours ago
The LLM didn't send fake invoices, it didn't send spam, a person did. And, the tool they used to do it was an LLM. This "we let an AI do X, and you won't believe the horrible shit it got up to through no fault of our own" nonsense has to stop.
Comment by piterrro 19 hours ago
Make - was never described how, Im actually surprised LLM didnt plan to print money. As much money - what does it mean? How much is much? As possible - there is no flavour of time, effort, cost and profit for the LLM. Could be even infinite, the result would be the same.
Given that the above goal is closest to „use cheating or unethical actions to create a profit” - I think the authors of it actually expected LLM to go wild.
Also
> Going forward, we plan to recreate this experiment with longer time horizons but using simulated environments instead.
Watch out, they will try that again.
Comment by sdeframond 19 hours ago
Comment by aitchnyu 19 hours ago
Comment by matthewiiiv 19 hours ago
a ceo agent spins up a bunch of AAARRR sub agents each morning and they pitch an idea to implement. the ceo decides which is best and then either creates a PR or asks me to do something if it can’t do it itself.
Comment by matkoniecz 18 hours ago
Comment by fatata123 1 hour ago
Comment by pydry 19 hours ago
Comment by redox99 19 hours ago
Comment by majorbugger 19 hours ago
Comment by wesleywt 19 hours ago
Comment by almost 19 hours ago
Comment by pluc 19 hours ago
Comment by ronsor 19 hours ago
"Computer says no."
Comment by hnlmorg 19 hours ago
Maybe I should create an agent that robs a bank. Or worse yet, pirates a movie.
Comment by altmanaltman 19 hours ago
They should be banned from posting on here after that. Like you are willingly spamming people looking for work. And seem to be neutral about all harm called. The victim wrote to us "STOP", very interesting. such evil agents
Comment by Computer0 19 hours ago
Comment by AngryData 19 hours ago
Comment by altmanaltman 19 hours ago
Comment by Quekid5 19 hours ago
Comment by DataDaemon 19 hours ago
Comment by prymitive 19 hours ago
Comment by threecheese 19 hours ago
Comment by pseudosavant 19 hours ago
Comment by goatlover 19 hours ago
Comment by well_ackshually 19 hours ago
Keep moving the goalposts
Comment by rvz 19 hours ago
It's Artificially Generated Invoices.
Comment by ghostly_s 19 hours ago
Comment by rtk59931k 2 hours ago
Comment by sheetlite_74536 17 hours ago
Comment by robotswantdata 19 hours ago
Every time a tech lab unleashes an AI agent that breaks the law, this community treats it like a quirky engineering edge case. "Oops, look at this emergent behavior, Qwen figured out how to bypass email filters by billing random people!" No, it didn't figure out a clever hack. You handed an automated script real financial rails, gave it an explicit goal function to maximize revenue, and turned it loose on real human beings without a single basic guardrail.
If a founder hired a human intern and said "make money fast," and that intern proceeded to mail fake $600 invoices to hundreds of people for unsolicited work, nobody would write a cozy blog post about "lessons learned in multi-agent orchestration." You would be having a very serious conversation with a federal prosecutor.
Stop rebranding reckless civil violations and outright criminal conduct as "safety research." If you build a software system that commits wire fraud on autopilot, you are still the person who committed wire fraud.
Comment by htl 14 hours ago
Comment by miguelfrreg 19 hours ago
Comment by ElProlactin 19 hours ago
Comment by voidnullvalue 19 hours ago
Comment by ElProlactin 19 hours ago
Comment by javcasas 19 hours ago
Comment by Tade0 19 hours ago
I am an extremely naive person in all aspects of life, except for money.
I was genuinely with you until I started reading that last paragraph.
Comment by ElProlactin 19 hours ago
And that's why you're not making $400,000+/month passively. Only a select few have the ambition and drive to be a part of this. That's OK.
Comment by c22 18 hours ago
Comment by binarymax 19 hours ago
I mean it’s obviously not true - since if it were then why would you waste your time with a masterclass?
So the sarcasm bit is I don’t know if you’re parroting all the other masterclass scammers out there for lols or if you actually are a masterclass scammer.
Comment by Quekid5 19 hours ago
Comment by meindnoch 19 hours ago
Comment by ElProlactin 19 hours ago
Comment by operatingthetan 19 hours ago