HUSTLE

CNN Sues Perplexity AI: What Every Founder Needs to Know

CNN Perplexity AI copyright lawsuit implications for founders and startups
0:00
0:00🎧 21 min

In October 2025, CNN and Perplexity AI sat across the table and tried to make a deal. CNN would license its reporting to Perplexity’s AI search engine. Perplexity’s paid subscribers would get access to CNN’s paywalled stories. The two sides haggled over terms for weeks. By November, the talks collapsed. What happened next is the part every AI founder should pay attention to: Perplexity kept scraping CNN’s content anyway.

On May 28, 2026, CNN filed a 54-page federal lawsuit in the Southern District of New York accusing Perplexity of copying more than 17,000 stories, photos, and videos without permission. The complaint alleges Perplexity bypassed CNN’s technical blockers, used the content in real time to generate AI search responses, and then falsely told its “Comet Plus” subscribers they were getting premium CNN access. No such deal existed.

The CNN lawsuit against Perplexity AI is the latest in a pile-on that now includes nine separate copyright suits from the New York Times, News Corp, Dow Jones, Reddit, Encyclopedia Britannica, and others. For Perplexity, a company valued at $21 billion with $200 million in annual recurring revenue, these lawsuits represent an existential question about its business model. For every other AI founder building products that touch third-party content, this case is a preview of what’s coming.

Last updated: June 2026

Quick answers

What does the CNN Perplexity lawsuit mean for AI startups?

The CNN v. Perplexity lawsuit signals that publishers will sue AI companies that scrape content without licensing deals, even after failed negotiations. For founders, it means any AI product that retrieves, summarizes, or redistributes third-party content in real time faces growing legal exposure. The safe move is building licensing infrastructure before you scale, not after a complaint lands.

Can AI companies legally scrape web content?

It depends on the use. The 1991 Feist ruling established that raw facts aren’t copyrightable, but original creative expression is. Courts haven’t settled whether AI-powered retrieval and summarization counts as fair use. CNN’s complaint specifically targets real-time redistribution of original reporting, which is legally distinct from scraping data for model training. The case law is still forming.

Is Perplexity AI in legal trouble?

Yes, significantly. Perplexity faces nine active copyright lawsuits from CNN, the New York Times, News Corp, Dow Jones, the Chicago Tribune, Reddit, Encyclopedia Britannica, Merriam-Webster, and Japan’s Yomiuri Shimbun. The company has filed motions to dismiss the NYT and Chicago Tribune suits, arguing users, not Perplexity, cause any infringement. No court has ruled on the merits yet.

What CNN is actually alleging

CNN’s complaint makes three distinct legal claims, and founders building AI products need to understand each one because they represent different categories of risk.

The first claim is straightforward copyright infringement. CNN says Perplexity copied 17,000+ stories, photos, and videos from CNN’s platforms and third-party hosts, then used that content as real-time input to generate AI responses. This isn’t about training data. CNN alleges Perplexity’s system pulls their reporting live, processes it through its large language models, and serves up responses that compete directly with CNN for audience attention and ad revenue.

The second claim involves circumventing technical protections. CNN says it specifically blocked Perplexity’s web crawlers using robots.txt and other technical measures. Perplexity allegedly bypassed those blocks. Under the Digital Millennium Copyright Act, deliberately circumventing access controls is a separate violation from the copying itself.

The third claim is trademark-related and potentially the most damaging for Perplexity’s reputation. CNN alleges that when users asked Perplexity’s chatbot “What is Comet Plus?”, it responded by falsely claiming CNN was part of a “premium news bundle” similar to Apple News Plus. According to CNN’s complaint, no such partnership exists. This goes beyond scraping into false advertising territory.

Perplexity’s response was four words: “You can’t copyright facts.” Chief communications officer Jesse Dwyer positioned it as a First Amendment issue. That framing will be tested in court.

AI startup facing copyright lawsuit legal documents on desk

Why “you can’t copyright facts” won’t be enough

Perplexity’s defense leans on Feist Publications v. Rural Telephone Service, the 1991 Supreme Court case that established facts themselves aren’t copyrightable. In that case, a phone directory publisher copied another company’s listings. The Court ruled that an alphabetical list of names and numbers lacked the originality required for copyright protection.

The problem for Perplexity is that CNN’s content isn’t a phone directory. It’s original reporting with creative expression: investigative journalism, feature writing, video production, editorial judgment about what to cover and how to frame it. The Feist ruling protects the right to state that “inflation hit 3.2% in May.” It doesn’t protect the right to copy CNN’s 2,000-word analysis of what that number means for the housing market.

There’s another wrinkle that makes Perplexity’s position weaker than OpenAI’s in its fight with the New York Times. OpenAI’s core argument is that training an AI model on copyrighted text is transformative use. The model learns patterns from millions of documents and generates new text. Perplexity’s product does something different: it retrieves specific content in real time and uses it to generate responses that directly answer the same questions the original article answers. That’s closer to redistribution than transformation.

“The distinction between training-time use and inference-time retrieval is the single most important legal question in AI right now,” according to Traverse Legal’s analysis of AI copyright litigation. Training is a one-time process that creates something new. Real-time retrieval creates a competing product from someone else’s work, every single query.

How does this compare to the NYT vs. OpenAI lawsuit?

The NYT sued OpenAI in December 2023, alleging its articles were used to train GPT models without a license. As of mid-2026, that case is in the settlement discussion phase, with courts pushing both sides toward mediation. OpenAI now faces more than 30 active lawsuits, but its core legal argument is different from what Perplexity has to work with.

OpenAI argues its models transform copyrighted inputs into genuinely new outputs. The company doesn’t retrieve specific NYT articles to answer user queries. GPT-4 learned from those articles during training, but the training happened once. The model itself is the product, not the articles it was trained on.

Perplexity’s product works differently. Its “answer engine” fetches content from publisher websites in real time, processes it through AI models, and delivers summarized responses with citations. This is the same dynamic driving the zero-click search trend reshaping how users interact with content in 2026. When someone asks Perplexity about a breaking news story, the system is pulling from the original reporting right now, not recalling patterns from a training process that happened months ago.

This distinction matters because fair use analysis weighs whether the new work replaces the market for the original. A chatbot that retrieves and summarizes a CNN article about a Senate hearing is replacing the reason someone would visit CNN.com. A language model that learned writing patterns from millions of articles and generates an original response about Senate procedures is doing something qualitatively different.

For founders, the practical takeaway: if your AI product retrieves and repackages specific content to answer specific queries, you’re in Perplexity’s legal category. If it generates responses from trained knowledge without real-time retrieval, you’re in OpenAI’s category. Neither is fully resolved in court, but the retrieval model carries higher risk.

The two-track publisher response that’s splitting the industry

Not every publisher is suing. The media industry has split into two camps, and the divide tells founders exactly where the market is heading.

On one side: CNN, the New York Times, News Corp, Reddit, and Encyclopedia Britannica are litigating. They want courts to establish that AI companies must pay for content, and they’re willing to spend years in court to set that precedent.

On the other side: Time, Gannett, Le Monde, and Der Spiegel signed licensing deals with Perplexity. Le Monde’s arrangement is particularly telling. The French publisher signed in May 2025 and now shares 25% of the licensing revenue with its staff journalists, according to Press Gazette. Le Monde’s CEO has publicly urged other publishers to sign AI partnerships rather than fight them.

OpenAI took the licensing-first approach after the NYT lawsuit forced its hand. The Associated Press signed a deal with OpenAI in July 2023, becoming the first major publisher to do so. Google followed with its own AP licensing deal in 2025. By early 2026, Reach (a major UK publisher) signed with Amazon for Nova AI content access, and News Corp struck a deal with Meta.

The pattern is unmistakable: the AI companies that build licensing infrastructure survive. The ones that don’t build it get sued until they either build it or shut down the product.

What should AI founders do to avoid Perplexity’s legal exposure?

Four specific actions every AI founder should take before their product faces a copyright complaint.

1. Audit your data pipeline today. Map every source your AI product touches at inference time. If your system retrieves content from publisher websites, news APIs, or third-party databases to generate responses, you’re in the high-risk category. Boyer Law Firm’s 2026 compliance guide for AI startups recommends creating a “training-data register” that tags source rights for every content category your models access. California’s AB 2013, which took effect January 1, 2026, already requires companies to disclose dataset sources publicly.

2. Respect robots.txt as a legal document, not a suggestion. CNN’s complaint specifically alleges that Perplexity bypassed technical access controls. Under the DMCA, circumventing those controls is a separate legal violation. Several legal analyses of web scraping in 2026 note that courts increasingly treat robots.txt blocks as expressions of the site owner’s access restrictions. Ignoring them strengthens a plaintiff’s case.

3. Build licensing infrastructure before you need it. The cost of a content licensing deal is a fraction of the cost of defending a federal copyright lawsuit. Perplexity’s situation proves that failed negotiations followed by continued scraping creates the worst possible legal position. If you can’t afford to license content from a publisher, you can’t afford to build a product that depends on their content.

4. Separate your training pipeline from your inference pipeline. The legal risk profile is different for each. Training on publicly available text for the purpose of building a general-purpose model has a stronger fair use argument than retrieving specific articles in real time to answer specific queries. If your product can function without real-time content retrieval, architect it that way. If it can’t, licensing is your only sustainable path.

AI startup team reviewing content licensing strategy on laptop screen

The consolidation risk founders aren’t talking about

Here’s the scenario that should worry every AI startup founder more than a lawsuit: if the legal landscape settles on mandatory licensing for any AI product that touches publisher content, only the biggest companies can afford the infrastructure.

Google already has licensing deals with major publishers through Google News Showcase, a program it launched in 2020. OpenAI has been signing publisher deals since mid-2023, and is now offering $2M in tokens to YC startups to lock in the next generation of AI builders. Anthropic is reportedly in discussions with multiple media groups. These companies can absorb the cost of licensing because their revenue supports it. Google’s parent Alphabet generated $350 billion in revenue in 2025. OpenAI is targeting a $1 trillion IPO valuation.

Perplexity, at $21 billion and roughly $200 million in ARR, sits in the dangerous middle. Big enough to get sued, not yet big enough to outspend the litigation. Perplexity hit $500 million in annualized revenue by April 2026 and serves over 100 million monthly active users. But nine simultaneous copyright suits from some of the world’s largest media companies will burn through legal budgets fast.

For smaller AI startups, the math is worse. If licensing becomes a prerequisite for any product that summarizes, aggregates, or synthesizes published content, the barrier to entry goes up dramatically. A two-person startup building an AI research assistant can’t negotiate content deals with 50 publishers. The solo founders building million-dollar AI businesses in 2026 are mostly selling tools and services, not aggregating third-party content. This is how copyright enforcement becomes a moat that protects incumbents.

The counterargument is that standardized licensing frameworks could emerge. If the current wave of publisher-AI deals establishes market rates, smaller companies could license content through aggregators or clearinghouses, similar to how ASCAP and BMI handle music licensing. That infrastructure doesn’t exist yet, but the demand for it is growing.

What the Perplexity lawsuits mean for AI search specifically

Perplexity isn’t the only AI search product at risk. Google’s AI Overviews, which summarize web content at the top of search results, face the same fundamental question: does an AI-generated summary that replaces the need to visit the source constitute infringement?

Google has approached this by cutting licensing deals preemptively and by arguing that its summaries drive traffic back to publishers. Perplexity’s model is harder to defend on the traffic argument because its entire product is designed to give users the answer without clicking through to the source.

When Perplexity CEO Aravind Srinivas described his vision of AI delivering “meaningful answers rather than just retrieving links,” he was describing a product that directly competes with the publishers whose content powers it. That tension was always going to end up in court.

For founders building AI search or research tools, the lesson from Perplexity’s nine lawsuits is that the “answer engine” model needs a content acquisition strategy from day one. You can’t build a product that replaces the need to visit publisher websites and then argue you’re helping those publishers. The economics don’t support it, and the courts won’t either.

The timeline that matters

No court has ruled on the merits of any AI copyright case involving real-time retrieval. The NYT v. OpenAI case is in mediation. Perplexity filed motions to dismiss the NYT and Chicago Tribune suits in February 2026, arguing that any infringement is caused by “atypical, litigation-driven user behavior,” not Perplexity’s system itself. Those motions haven’t been decided.

The CNN case, filed in the same Southern District of New York courthouse where several other Perplexity suits are pending, could be consolidated with existing litigation. That would create a mega-case covering copyright, trademark, and DMCA claims from multiple plaintiffs. A consolidated ruling would set broader precedent faster than individual cases.

Founders shouldn’t wait for court rulings to act. The legal trend is clear: publishers are increasingly willing to sue, and the arguments for mandatory licensing are getting stronger with each new complaint. OpenAI’s confidential IPO filing and its ongoing effort to settle copyright disputes suggest even the largest AI company in the world recognizes that the era of free content is ending.

The smart move for any AI founder is to treat content licensing as a cost of doing business, the same way SaaS companies treat cloud hosting. You can try to run your product on scrapped content without paying for it. Perplexity tried that. Now they’re defending nine lawsuits while trying to grow a business valued at $21 billion.

Or you can build licensing into your cost structure from the start, the way founders building AI-native companies in 2026 are learning to budget for compliance alongside compute. The startups that treat content like infrastructure will outlast the ones that treat it like something free on the internet.

Read More From the HUSTLE desk