AI, law degrees, and the real work in front of legal teams
Jad Tarifi, former head of Google's first generative AI team, argues that law degrees are losing relevance as AI makes rote memorization obsolete. It's a sharp take, and it is pushing a real debate. But the future he sketches is not here yet. For now, the most valuable lawyers are the ones who convert AI uncertainty into clear policy, defensible risk, and enforceable deals.
Publishers vs. AI companies: the Perplexity disputes
Lawsuits continue to test how copyrighted content can be used to train and run large language models. A New York federal court refused to dismiss or transfer a case brought by News Corp, Dow Jones, and the New York Post against Perplexity AI over alleged misuse of their articles. A major Japanese newspaper company has also filed suit. The BBC has threatened legal action, and Forbes and WIRED have publicly criticized Perplexity's use of their content.
This week, Encyclopedia Britannica and Merriam-Webster filed a new case alleging "massive copying," including the use of undisclosed crawlers that evade site protections and pull full articles into an AI database without permission. The complaint claims Perplexity's responses are near-verbatim to Britannica content and that the service diverts traffic by summarizing articles while generating few click-throughs from citations. Beyond copyright, the filings raise trademark issues tied to AI "hallucinations" that attribute false statements to trusted brands. Britannica also argues Perplexity included its copyrighted works in a RAG system that connects models to external sources.
Why this matters for legal teams
Courts are being asked whether training and answer-generation that substitutes for reading the original work is permissible. Product leaders see the same warning: innovate with AI while staying inside copyright, contract, and trademark boundaries. Legal teams need concrete controls now, not just position papers.
Core legal issues to track
- Copyright infringement (training and outputs): Does model training on copyrighted text qualify as fair use? Do near-verbatim answers tip the analysis against fair use, especially when they displace visits to the source?
- Anti-circumvention (DMCA §1201): Allegations of "stealth" crawlers that bypass technical measures could trigger DMCA claims beyond basic infringement.
- Contract and access claims: Breach of terms of service, trespass to chattels, and Computer Fraud and Abuse Act theories may come into play if site restrictions or authentication barriers are bypassed.
- Trademark and false attribution: If the system outputs incorrect statements and ties them to brands like Britannica or Merriam-Webster, plaintiffs may pursue false association or false advertising claims.
- Market harm and remedies: Publishers allege diverted traffic and lost revenue due to summaries that replace clicks. Expect arguments for damages and potential injunctions, including model or dataset deletion orders.
Action checklist for in-house and outside counsel
- Data sourcing audit: Map training, fine-tuning, and RAG sources. Record licenses, permissions, robots.txt compliance, and any scraping settings or user agents.
- RAG governance: Use allowlists/denylists, enforce cache time-to-live, and honor no-scrape/no-index flags. Log retrieval provenance so you can trace what fed a specific answer.
- Licensing strategy: Where content is core to product value, secure licenses or APIs with clear use rights, attribution requirements, and rate limits.
- Output controls: Detect and block near-verbatim reproduction. Add source attribution, link prominence, and citation placement that actually earns clicks.
- Brand and hallucination risk: Filter for brand mentions, require fact sources for brand claims, and add disclaimers where appropriate. Set a takedown and correction protocol.
- Supplier and indemnity terms: Tighten contracts with model vendors, data providers, and tooling partners. Seek indemnities, audit rights, and cooperation clauses for legal holds.
- Incident readiness: Prepare playbooks for content owner complaints, DMCA notices, and emergency model or index changes. Preserve logs for litigation.
- Board and insurance alignment: Brief leadership on exposure ranges (statutory vs. actual damages, injunction risk) and confirm coverage positions with carriers.
Open questions to monitor
- How courts apply fair use to both model training and answer-generation that may substitute for reading the original work.
- Whether alleged bypassing of technical measures sticks under DMCA §1201 and how "protective measures" are defined.
- If courts treat retrieval-augmented generation differently from pretraining with respect to liability and remedies.
- Remedial scope: licensing, damages, deletion of datasets, or restrictions on future use.
- Venue and procedural skirmishes that affect timing and leverage.
What this means for legal talent
Degrees are not irrelevant. What changes is the value curve: less credit for memorizing doctrine, more credit for shaping policy, deal structures, and controls that keep AI products shippable and defensible. The lawyers who win here combine IP fluency, product sense, and operational discipline.
Further reading
Skill up
If your team needs a structured primer on AI concepts and workflows to inform policy and contracts, see Complete AI Training: Courses by Job.
