#AI Training Data

Every story tagged AI Training Data, curated for CIOs and IT leaders — ranked by source credibility, engagement, and freshness.

44 stories · open in the command center

  • AI & MLHacker News3m

    The Scientific Literature Is Poisonous to LLMs

    Scientific literature contains significant quality issues—including half-truths, convenient omissions, and outright fraud—that actively degrade LLM performance, with empirical research showing that removing major academic corpora (ArXiv, PhilPapers, NIH ExPorter) actually improves AI model accuracy and reduces toxic outputs. This finding has critical implications for AI-for-science initiatives, as AI systems cannot access the informal social networks and reputational verification that human scientists use to filter reliable information, creating a fundamental trust and validation gap. Organizations leveraging AI for research and scientific decision-making must recognize that training data quality from academic sources is unreliable and develop alternative verification mechanisms rather than assuming published literature provides trustworthy foundations for AI applications.

  • AI & MLHacker News3m

    AI companies are shredding rare books

    AI companies are systematically purchasing and destroying rare books through industrial scanning operations to train large language models, with a federal court ruling this practice legal as "fair use." This creates irreversible loss of irreplaceable historical knowledge while enabling vendors to obscure the practice through NDAs and euphemistic marketing, establishing a precedent that will likely accelerate similar data acquisition strategies across the industry. For IT organizations, this signals emerging legal and reputational risks around AI training data sourcing, supply chain opacity, and the need to establish ethical data governance policies before regulatory backlash forces compliance.

  • Security & PrivacyTechMemeMunsif Vengattil2m

    Delhi High Court rules OpenAI's use of ANI content to train ChatGPT wasn't copyright infringement as the news agency didn't show ChatGPT reproduced its reports (Munsif Vengattil/Reuters)

    A Delhi High Court ruling establishes that using copyrighted content for AI model training may not constitute copyright infringement if the trained model doesn't directly reproduce the original works, creating significant legal precedent that reduces liability exposure for organizations deploying generative AI systems. This decision signals that fair use protections extend to AI training practices in India and globally, potentially enabling broader use of proprietary data for AI development without explicit licensing agreements. Technology leaders should recognize this as a favorable regulatory tailwind for AI initiatives, though the ruling's applicability to other jurisdictions remains uncertain and may invite legislative responses.

  • AI & MLTechMemeCJ Haddad2m

    Reddit's stock closed down 8.32% after a report that the company was considering ending Google's access to its content for AI training; RDDT is down ~26% YTD (CJ Haddad/CNBC)

    Reddit's consideration of restricting Google's AI training access to its content triggered an 8.32% stock decline, signaling market concerns about monetization strategy and AI partnership value. This move reflects broader tensions between content platforms and AI companies over data usage rights and revenue sharing, presenting IT leaders with critical decisions about data governance, AI training dependencies, and competitive positioning. Organizations should reassess their own content licensing agreements and AI training data sourcing strategies, as this trend could reshape how enterprises access and leverage third-party data for AI initiatives.

  • AI & MLHacker News3m

    Judge approves $1.5B Anthropic settlement for pirated books used to train Claude

    A judge has approved a $1.5 billion settlement for Anthropic over the use of pirated books in training its Claude AI model, establishing significant financial and legal precedent for generative AI development. This settlement highlights critical intellectual property and data governance risks that IT leaders must address in their AI initiatives, as unauthorized use of copyrighted content can result in substantial liability and regulatory exposure. Organizations adopting or developing AI systems must now implement robust data provenance controls, licensing verification, and compliance frameworks to avoid similar costly litigation.

  • Startups & FundingTechMemeÉanna Kelly2m

    Munich-based Microagi, which collects factory and household data to train humanoid robots, raised $55M led by Hummingbird in Germany's largest ever seed round (Éanna Kelly/Sifted)

    Microagi's $55M seed round represents a significant market validation for humanoid robotics trained on real-world factory and household data, signaling that enterprises should prepare for widespread robotic automation across operations and facilities management. This breakthrough funding signals accelerating convergence of AI, robotics, and operational technology, requiring IT leaders to evaluate workforce planning, data governance, and infrastructure readiness for human-robot collaboration at scale. German venture capital dominance in this space suggests that organizations must monitor European tech innovation clusters and consider strategic partnerships or capabilities development in humanoid robotics to maintain competitive advantage.

  • Security & PrivacyThe VergeJess Weatherbed2m

    Suno snatched millions of songs from YouTube, Genius, and Deezer

    Suno's AI music generator was trained on millions of copyrighted songs scraped from YouTube Music, Deezer, Genius, and other platforms without authorization, raising critical legal and compliance risks for organizations adopting AI-generated content tools. This breach exposes broader vulnerabilities in AI supply chain security and intellectual property protection, requiring IT leaders to reassess vendor due diligence, data governance, and legal liability exposure when deploying third-party AI solutions. The incident underscores the urgent need for enterprises to establish policies around AI model provenance and to prepare for potential regulatory crackdowns on data scraping practices.

  • Security & PrivacyTechCrunchAmanda Silberling2m

    Hack suggests AI music generator Suno scraped YouTube for training data

    AI music generator Suno was compromised via supply chain attack, exposing source code that allegedly shows unauthorized scraping of copyrighted content from YouTube and other platforms, alongside customer data including email addresses and partial credit card numbers. This incident underscores significant legal and compliance risks for organizations adopting AI tools, as major record labels are actively pursuing DMCA violations against Suno and similar platforms that circumvent platform protections to train models on copyrighted material. IT leaders must now evaluate the intellectual property, data security, and regulatory exposure associated with third-party AI vendors, particularly regarding how training data is sourced and whether vendors maintain adequate security controls.

  • Security & PrivacyTechCrunchSarah Perez2m

    If you use Google, you’re training its AI. Here’s how to opt out.

    Google has quietly expanded its AI training practices to include user-uploaded media (images, audio, video files) across its services through a restructured privacy settings update, with data now being used to train AI models unless explicitly opted out. This reflects a broader industry trend of leveraging user-generated content for AI improvement, creating significant data governance and compliance risks for enterprises whose employees use Google services. IT leaders must understand that their organization's sensitive data—from Google Lens images to voice recordings—may be retained for AI training purposes, necessitating immediate policy review and potential restrictions on consumer Google service usage.

  • Startups & FundingTechCrunchSean O'Kane2m

    Robot hand company settles Tesla trade secret suit and announces $11M raise

    Proception Robotics, founded by a former Tesla Optimus engineer, has resolved its trade secret litigation and secured $11M in funding to commercialize advanced robotic hands with 22 degrees of freedom—addressing what industry leaders identify as a critical bottleneck in humanoid robot deployment. The company's differentiated approach combines sophisticated hardware with scalable, sensor-based data collection methods that could accelerate practical dexterous manipulation capabilities by years compared to industry consensus timelines. For IT and technology organizations, this signals that specialized robotics components are moving from internal R&D to third-party suppliers, creating both integration opportunities and dependencies that require strategic evaluation.

  • Security & PrivacyWiredReece Rogers2m

    How to Opt Out of Google Search’s New AI Data Training Feature

    Google is automatically enrolling users' search data—including images, audio, and video—into AI model training via a new default-enabled Search Services History feature, requiring manual opt-out through account settings. This represents a significant expansion of data collection practices across Google's integrated services ecosystem, with trained data persisting for up to 4 years even after deletion, creating potential privacy and compliance risks for enterprise users and their organizations. IT leaders must evaluate the implications for employee data privacy policies, contractual obligations with Google, and the growing pattern of tech vendors shifting from opt-in to opt-out models for AI training, which may necessitate updated corporate technology governance frameworks.

  • AI & MLThe VergeTerrence O’Brien2m

    The Atlantic created a searchable database of the music used to train AI

    The Atlantic discovered millions of copyrighted music tracks in publicly accessible AI training datasets, highlighting a critical intellectual property and compliance risk for organizations developing generative AI models. This exposure reveals that major tech companies may be inadvertently—or deliberately—violating content licensing agreements and platform terms of service, creating potential legal liability and regulatory scrutiny for enterprises deploying AI systems. IT leaders must now audit their AI training data sources and governance practices to mitigate IP infringement risks and ensure compliance with emerging AI regulations.

  • AI & MLTechCrunchTim Fernholz2m

    Collecting robot training data is dirty, unglamorous work. Some AI labs are already paying XDOF to do it.

    Major AI labs are racing to develop robotics capabilities but face a critical bottleneck: the lack of high-quality training data for physical AI systems. XDOF, a newly launched startup backed by $70M in venture funding, is positioning itself as the essential infrastructure provider for robot training data collection, annotation, and pipeline management—addressing a market gap that even frontier AI labs find too operationally complex to build themselves. This emerging data infrastructure business represents a strategic dependency that IT leaders and CIOs should monitor closely, as it mirrors the critical role data infrastructure played in the LLM race and will likely become a key competitive lever for AI-driven automation initiatives.

  • Security & PrivacyTechMemeErnesto Van der Sar2m

    A US judge rejects Meta's bid to dismiss a lawsuit from Strike 3, which owns porn production companies, that alleges Meta torrented its videos for AI training (Ernesto Van der Sar/TorrentFreak)

    A federal judge has allowed Strike 3 Holdings' lawsuit against Meta to proceed, ruling that the company need not prove Meta specifically used their copyrighted content for AI training—only that unauthorized torrenting occurred. This decision significantly increases Meta's legal liability exposure for large-scale data collection practices and establishes concerning precedent for how AI training datasets are legally scrutinized, with potential implications for other tech companies' AI development pipelines. IT and data governance leaders must reassess their organization's content sourcing, licensing compliance, and AI training data provenance practices to mitigate similar legal and reputational risks.

  • AI & MLThe VergeTerrence O’Brien2m

    Google won’t just admit it’s feeding YouTube creators to its music AI

    Google is using YouTube creator content to train its Lyria music AI model but refuses to explicitly confirm this practice, instead relying on broad terms-of-service language and legal ambiguity to maintain plausible deniability amid ongoing litigation from independent musicians. This situation highlights critical risks for IT organizations managing AI training data governance, content licensing compliance, and creator relationships—establishing clear data usage policies and transparent AI training practices is now essential to mitigate legal exposure and maintain stakeholder trust. Technology leaders must recognize that leveraging user-generated content for AI training without explicit consent creates escalating regulatory and reputational risks that extend beyond individual litigation.

  • Security & PrivacyThe VergeEmma Roth2m

    Google will save your Lens photos, Search Live recordings, and Translate audio for AI training

    Google is expanding its data collection practices by automatically saving user interactions across Google Lens, Search Live, voice searches, and Translate for AI model training and personalization, creating significant privacy and compliance implications for enterprises managing corporate data use. This change separates from existing privacy controls, requiring IT organizations to reassess data governance policies and employee awareness programs, particularly regarding how user-generated content feeds AI development. CIOs must understand the new 'Search Services History' setting and its business implications, as this represents a broader industry trend of leveraging user interactions for AI advancement that may conflict with organizational data retention and privacy policies.

  • AI & MLTechMemeJason Koebler2m

    Moderators of the r/biohackers subreddit claim that peptide companies are spamming their forum in hopes of getting their posts scraped and used by AI chatbots (Jason Koebler/404 Media)

    Bad actors are systematically poisoning public data sources like Reddit to manipulate AI model training and search results, creating a new attack vector that threatens data integrity and AI reliability across enterprise systems. This highlights critical risks in your organization's AI supply chain, as models trained on compromised public data could propagate misinformation or generate unreliable outputs for business-critical decisions. IT leaders must implement data validation, source authentication, and AI governance frameworks to protect against training data contamination and ensure the trustworthiness of AI systems your organization depends on.

  • AI & MLTechMemeJason Koebler2m

    Source: Google asked Google Play app devs to join a "confidential content offer pilot", offering to pay for access to codebases to use them to train AI tools (Jason Koebler/404 Media)

    Google is directly acquiring source code from Android developers through a confidential pilot program to train its AI models, signaling a strategic shift toward securing proprietary training data and potentially creating competitive advantages in AI-powered development tools. This practice raises significant intellectual property, security, and competitive concerns for IT organizations, as it establishes a precedent where third-party code becomes a commodity for AI training and may influence future vendor relationships and code governance policies. Technology leaders should prepare for similar acquisition programs from other cloud and AI vendors, which could reshape how organizations value, protect, and monetize their proprietary codebases.

  • Startups & FundingTechMemeBen Weiss2m

    New York City-based Mecka AI, which trains robots with human data sourced from body sensors and iPhones, raised $60M, including a $25M Series A (Ben Weiss/Fortune)

    Mecka AI's $60M funding round signals a critical shift in robotics development where human motion data becomes a strategic competitive asset, requiring IT organizations to evaluate their data infrastructure capabilities for capturing, securing, and managing biometric sensor data at scale. For enterprises, this advancement means robotics automation is accelerating toward practical deployment, necessitating IT teams to prepare infrastructure, security protocols, and talent strategies to support robot integration across operations. The reliance on consumer devices (iPhones) and body sensors to train AI models also raises important data governance and privacy considerations that IT leaders must address proactively.

  • Security & PrivacyTechMemeIvan Mehta2m

    Strava is adding a $11.99 monthly fee for developer API access and moving public profiles and fitness club listings behind authentication to combat AI scraping (Ivan Mehta/TechCrunch)

    Strava is implementing monetization and access controls for its developer API ($11.99/month) while restricting public data access to combat unauthorized AI training data scraping, signaling a broader industry shift toward protecting proprietary data assets. This move reflects growing tension between AI companies' data acquisition needs and platform providers' IP protection, forcing IT organizations to reconsider API strategies, data governance, and third-party integration dependencies. Technology leaders should anticipate similar access restrictions and monetization models across other platforms as companies prioritize data control over open accessibility.

  • AI & MLThe VergeRobert Hart2m

    Tech companies desperately want to film you doing chores

    Tech startups are aggressively collecting real-world video data of people performing physical tasks (cleaning, cooking, laundry) to train robotics and physical AI systems, with some companies offering free services or direct payments in exchange for footage. This represents a significant shift in data collection strategies—moving from easily scraped digital content to monetized physical-world data—creating new privacy risks and ethical considerations that IT leaders must address in their organizations' policies and vendor management. CIOs should anticipate increased regulatory scrutiny around biometric and activity data collection, potential employee concerns about workplace monitoring, and the need for robust data governance frameworks as AI training becomes a primary business driver.

  • Startups & FundingArs TechnicaJeremy Hsu2m

    Startup offers free home cleaning—if it can record it all for robot training

    MicroAGI is leveraging a free home cleaning service to collect first-person video data for AI robotics training, representing an emerging business model where companies monetize training data collection through consumer incentives and crowdsourced recording. This trend signals that AI development costs are shifting toward data acquisition rather than traditional infrastructure, creating both opportunities for companies to reduce training expenses and significant risks around data privacy, consent management, and liability that IT organizations must monitor. Technology leaders should expect increasing pressure from business units to pursue similar data-collection strategies while preparing for potential regulatory, security, and reputational risks associated with personal data handling at scale.

  • Startups & FundingThe VergeRobert Hart2m

    This AI startup will clean your home for free to train future robots

    Shift, an AI startup, is offering free home cleaning services in exchange for video footage of cleaners performing tasks, which will be used to train robotic systems—representing a novel data acquisition model where consumers subsidize AI development through privacy exchange. This trend signals a strategic shift in how AI companies source training data and raises important governance questions for IT leaders around data privacy, vendor risk management, and the emerging practice of monetizing human activity recordings. Organizations should prepare for similar data-for-services models to proliferate across enterprise operations, requiring updated policies around third-party data collection and AI training data provenance.

  • Startups & FundingTechMemeIvan Mehta2m

    Human Archive, which trains robots using first-person video from 1,000+ camera-equipped caps worn by Indian home services workers, raised $8.2M from YC and more (Ivan Mehta/TechCrunch)

    Human Archive has secured $8.2M in funding to develop a novel approach to robot training using first-person video data collected from camera-equipped caps worn by home services workers in India, representing a significant shift toward practical, real-world data collection for AI/robotics development. This model demonstrates how enterprises can leverage distributed workforce data to accelerate automation capabilities, with implications for supply chain optimization, labor cost reduction, and operational efficiency across service industries. IT organizations should recognize this as an emerging pattern where edge data collection infrastructure becomes critical to competitive advantage in automation initiatives.

  • Startups & FundingTechCrunchIvan Mehta2m

    This startup is betting India’s gig economy can train the world’s robots

    Human Archive is capturing real-world task data from India's gig economy workers using specialized hardware (cameras, motion capture, tactile sensors) to train AI robots at scale—addressing a critical bottleneck in robotics development. This emerging data-as-a-service model, validated by $8.2M in funding from top-tier investors, represents a new strategic intersection of labor arbitrage, AI training infrastructure, and robotics, with implications for how enterprises source AI training data and manage the ethical dimensions of worker surveillance. IT leaders should anticipate increased demand for synthetic data pipelines, edge computing at scale, and new governance frameworks as robotics companies compete for high-quality training datasets.

  • AI & MLTechMemeCarmen Arroyo2m

    Internal chats: in March, xAI offered employees $420 in exchange for completed tax filings as training data for Grok, but the bonuses haven't been paid out (Carmen Arroyo/Bloomberg)

    Elon Musk's xAI attempted to use employee incentives ($420 bonuses) to collect personal tax data for training its Grok AI model, but failed to compensate participants, raising critical concerns about data governance, employee trust, and the legal/compliance risks of using personal financial information for AI training without proper safeguards. This incident highlights the tension between rapid AI development and responsible data handling practices, signaling that technology leaders must establish rigorous policies around consent, compensation, and data stewardship to avoid reputational damage and regulatory exposure.

  • Security & PrivacyArs TechnicaAshley Belanger2m

    Anthropic’s $1.5B copyright settlement is getting messy as judge delays approval

    A federal judge has delayed approval of Anthropic's $1.5 billion copyright settlement due to objections from authors challenging excessive attorney fees ($320M+ versus $3,000 per author payout) and inadequate protections against future AI training on pirated works. This decision creates significant legal and financial uncertainty for AI companies relying on large-scale content for model training, exposing potential gaps in settlement adequacy and class-action governance that could set precedent for future IP litigation. Technology leaders should anticipate prolonged regulatory scrutiny, higher compliance costs for AI training datasets, and potential contractual obligations to restrict future use of copyrighted materials.

  • Startups & FundingTechCrunchIvan Mehta2m

    Wirestock raises $23M to supply creative multi-modal data to AI labs

    Wirestock has pivoted from a stock photography platform to a high-growth AI data provider, raising $23M Series A to supply multi-modal creative datasets (images, videos, 3D content) to major foundation model makers, demonstrating the strategic value of curated, quality data in competitive AI development. With $40M ARR, 700,000 contributor artists, and contracts with six major AI labs, Wirestock exemplifies how enterprises can monetize existing data assets while the broader data supply market experiences explosive demand from hyperscalers racing to improve model capabilities. For IT organizations, this signals that data procurement, quality assurance, and contributor management are becoming critical enterprise competencies—and that specialized data supply partnerships may be essential infrastructure for competitive AI implementation.

  • Startups & FundingTechCrunchRussell Brandom2m

    Origin Lab raises $8M to help video game companies sell data to world-model builders

    Origin Lab has secured $8M in funding to create a marketplace connecting video game companies' digital assets with AI labs building world models for robotics and physical simulation, addressing a critical data bottleneck for next-generation AI systems. This represents a significant emerging market opportunity for data infrastructure vendors serving well-capitalized AI labs, similar to the success of Scale.AI, with implications for how enterprises source training data and manage AI development pipelines. IT leaders should recognize this as part of a broader shift toward specialized data supply chains that will become essential dependencies for deploying advanced AI capabilities.

  • Security & PrivacyThe VergeEmma Roth2m

    Meta sued by major book publishers over copyright infringement

    Meta faces a major class action lawsuit from five major publishers and an author alleging the company engaged in massive copyright infringement by training its Llama AI models on pirated books and journals without permission, with the model producing verbatim reproductions of copyrighted content. This lawsuit signals escalating legal and regulatory risks for IT organizations deploying generative AI, requiring immediate review of data sourcing, model training practices, and intellectual property compliance frameworks. The outcome could establish precedent for enterprise AI governance and force companies to implement stricter content licensing and attribution protocols in their AI development pipelines.

Browse all tags