OpenAI Reportedly Used More Than a Million Hours of YouTube Videos to Train Its Latest AI Model

Where does AI training data come from?

A report from The New York Times revealed on Friday that OpenAI may have trained AI models on YouTube video transcriptions and Google may have been doing the same thing.

The report found that in the hunt for fresh digital data to train its newer, smarter AI system, OpenAI researchers created a workaround called Whisper, which could take YouTube videos and transcribe them into text that could then be fed as new AI training data — for a more conversational, next-generation AI.

Get investing news alerts:

The process of developing GPT-4, the powerful AI model behind OpenAI's latest ChatGPT chatbot, took over a million hours of YouTube videos transcribed by Whisper, according to the NYTimes' sources.

The Times reports that OpenAI employees had conversations about how YouTube transcription training data could potentially violate YouTube's rules, but OpenAI decided to move forward anyway with the belief that training AI with the videos was fair use.

Knowledge of where the training data was coming from extended up to senior leadership, according to The Times, with OpenAI's president Greg Brockman even allegedly helping collect videos.

The Wall Street Journal's Joanna Stern interviewed OpenAI's CTO Mira Murati last month and asked her what data was used to train one of OpenAI's most recent products: a tool called Sora that generates videos based on text prompts.

"We used publicly available data and licensed data," Murati said. When Stern asked "So, videos on YouTube?" Murati replied, "I'm actually not sure about that."

When Stern further asked "Videos from Facebook, Instagram?" Murati stated, "You know, if they were publicly available, publicly available to use, there might be the data, but I'm not sure. I'm not confident about it."

YouTube CEO Neal Mohan said last week that if OpenAI used YouTube videos to train Sora, that would be a "clear violation" of YouTube's terms of use.

The terms of service "does not allow for things like transcripts or video bits to be downloaded," Mohan told Emily Chang, host of Bloomberg Originals.

Yet five sources told The Times that Google did the same thing as OpenAI, allegedly transcribing YouTube videos to generate new training text for its AI models in a potential violation of copyright law.

Google owns YouTube and told The Times that its AI is "trained on some YouTube content" that its agreements with creators allow.

Lawsuits over training AI with copyrighted material have become widespread in recent years, with authors like Paul Tremblay and Sarah Silverman alleging that their books were part of datasets used to train AI — without their consent.

The lawyers for these lawsuits, Joseph Saveri and Matthew Butterick, state on their website that generative AI is just "human intelligence, repackaged and divorced from its creators."

More than 15,000 authors signed a letter last year asking big tech CEOs, including ones at OpenAI, Google, Microsoft, Meta, and IBM, to obtain the consent of writers before training AI with their work and credit and compensate them.

It's not just authors: musicians too are feeling the impact of AI. Artists like Billie Eilish and Jon Bon Jovi signed an open letter last week accusing big tech companies of using their work to train models without permission or compensation.

"These efforts are direly aimed at replacing the work of human artists with massive quantities of AI-created "sounds" and "images" that substantially dilute the royalty pools that are paid out to artists," the letter stated.

Tennessee became the first state to pass legislation protecting artists from deepfakes, or cloned and manipulated versions of their voices, last month.

Where Should You Invest $1,000 Right Now?

Before you make your next trade, you'll want to hear this.

MarketBeat keeps track of Wall Street's top-rated and best performing research analysts and the stocks they recommend to their clients on a daily basis.

Our team has identified the five stocks that top analysts are quietly whispering to their clients to buy now before the broader market catches on... and none of the big name stocks were on the list.

They believe these five stocks are the five best companies for investors to buy now...

See The Five Stocks Here

Metaverse Stocks And Why You Can't Ignore Them

Thinking about investing in Meta, Roblox, or Unity? Enter your email to learn what streetwise investors need to know about the metaverse and public markets before making an investment.

Get This Free Report

Recent Videos

Stock Lists

All Stock Lists

Investing Tools

Calendars and Tools

Search Headlines

Get 30 Days of MarketBeat All Access for Free

Start Your 30-Day Trial

In-depth profiles and analysis for 20,000 public companies.
Real-time analyst ratings, insider transactions, earnings data, and more.
Our daily ratings and market update email newsletter.

Sign In
Create Account

Your Email Address:

Your Password:

Forgot your password?

Your Email Address:

Choose a Password:

By creating a free account, you agree to our terms of service. This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.

OpenAI Reportedly Used More Than a Million Hours of YouTube Videos to Train Its Latest AI Model

Where Should You Invest $1,000 Right Now?

Featured Articles and Offers

Recent Videos

Stock Lists

Investing Tools

Search Headlines

Best-in-Class Portfolio Monitoring

Stock Ideas and Recommendations

Advanced Stock Screeners and Research Tools

About MarketBeat

MarketBeat Products

Popular Tools

Financial Calendars

Terms & Info

OpenAI Reportedly Used More Than a Million Hours of YouTube Videos to Train Its Latest AI Model

Where Should You Invest $1,000 Right Now?

Featured Articles and Offers

Recent Videos

Stock Lists

Investing Tools

Search Headlines

MarketBeat All Access Features

Best-in-Class Portfolio Monitoring

Stock Ideas and Recommendations

Advanced Stock Screeners and Research Tools