Skip to content
Suno hack reveals scraped YouTube, Deezer, podcast training audio

Suno hack reveals scraped YouTube, Deezer, podcast training audio

Aiweekly.Co July 15, 2026

There is a specific kind of interesting when a security breach lands in the middle of an active copyright lawsuit, because it turns a lawyer's theory of the case into a document. That is what happened to Suno, the generative-music startup, according to 404 Media . A hacker going by the handle ellie.191 reportedly exploited the Shai-Hulud npm supply-chain worm to pull source code from 2023 and 2024 out of Suno, along with customer emails, phone numbers, and Stripe payment details, and then handed the material to reporters. The hacker told 404 Media they had 'no specific motivation for hacking Suno.'

The files spell out, in inventory form, where Suno's training audio came from. The reporting lists 2,013,545 clips from YouTube Music running to 113,879 hours, 12,287 hours from Deezer, 17,615 hours from Genius, 62,117 hours from Pond5, 3,726 hours from Jamendo, 19,514 hours from the International Music Score Library Project, and around a million hours of audio pulled from roughly 420,000 podcasts identified through RSS feeds. Code inside the leak reportedly used Bright Data, a commercial scraping infrastructure provider, to extract from YouTube, and included routines that specifically searched for acapella versions of songs.

The reason that matters is legal, not just embarrassing. The RIAA has been suing Suno for what it calls 'stream ripping' from YouTube, and Suno's own court filing already conceded its 'training data includes essentially all music files of reasonable quality that are accessible on the open internet.' A leaked inventory that names Deezer, Genius, Pond5, and YouTube by hour count moves that argument from RIAA allegation to Suno document. Suno's public position is still that training on copyrighted works is fair use.

The honest caveat is that Suno told 404 Media the leaked material is 'outdated source code that is no longer in use' and that 'no sensitive personal information was compromised,' while confirming it had been the subject of a 'limited security incident' discovered in November 2025 and declining to notify affected customers individually. What the reporting doesn't give you is whether the current training pipeline still uses the same sources at the same scale, or how a court will weigh leaked internal documents against the company's own filings.

The forward-looking part, for anyone building or investing around generative audio, is that the discovery phase in these cases just got a lot cheaper for plaintiffs. Every AI music and voice startup that scraped from consumer platforms should assume its ingestion inventory is one breach or one subpoena away from being on the record.

Coverage cluster as of 24h after publish

Adds a DMCA circumvention angle separate from copyright, notes the supply-chain entry in November 2025, and documents Suno's failure to notify affected customers.

Focuses on Suno's dismissive 'limited' public response against the documented scope of customer data exposed; frames Sony's fair-use claim as now evidentially stronger.

Surfaces the GEMA/Munich I Regional Court ruling due July 31, 2026, and distinguishes Warner's November 2025 licensing settlement from the still-live Universal and Sony suits.

Notes the breach landed six weeks after Suno's $400M raise at a $5.4B valuation and flags that customer data exposure was undisclosed to affected users at time of reporting.

Originally reported by 404media.co

Original headline: Suno Breached via Shai-Hulud npm Worm — Hacker 'ellie.191' Leaks Internal Files Showing Suno Scraped 2M+ YouTube Music Clips, 12,287 Hours of Deezer, 17,615 Hours From Genius, 62,117 Hours From Pond5 and ~1M Hours of Podcasts to Train Its Generative Music Models

Extracted Entities