top of page

The consent gap in India's physical AI gold rush

(An AI-generated charcoal sketch of women wearing cameras on their foreheads in India training AI on human skills like packing and folding for the future use in robotics.)



India is becoming one of the primary sources of training data for the physical AI revolution. Robots that fold laundry, cook, sort packages and navigate factory floors are being taught to move by watching humans do it first, and increasingly, those humans are Indian.


The scale of interest is real. Indian physical AI startups have raised roughly $155 million across 31 deals so far in 2026, up from $130 million in 2025, $124 million in 2024, and $91 million in 2023, a steady climb, not a spike (Entrackr, 2026). Globally, the category is moving even faster: physical AI funding worldwide crossed $47.4 billion in just the first half of 2026, nearly four times the total from the second half of 2025 (Crunchbase News, 2026).

Much of this is new money chasing a genuinely new kind of data. Unlike large language models, which could be trained on the open internet's text, embodied AI needs footage of real hands doing real physical work, in a kitchen, a factory line, a delivery route. That data mostly doesn't exist yet. It has to be collected, person by person, task by task. India, with its large workforce, low collection costs and deep outsourcing infrastructure, has become an obvious place to collect it.


New data companies are building exactly this pipeline. Some run head-mounted cameras across networks of workers to capture first-person video of everyday tasks: cooking, cleaning, delivery, factory work, at a scale of thousands of hours, sometimes from workers who may not fully understand what the footage becomes, who buys it, or how long it's kept.

That's the gap this article is about. In the rush to source physical-world data at scale, consent and permissions risk becoming a checkbox exercise rather than a real safeguard, and the companies buying this data may not realize how much they're exposed to.


Why consent isn't just an ethics problem. It's a balance-sheet problem


Two U.S. cases show what happens when the underlying consent fails.


In 2021, the US Federal Trade Commission (FTC), the US government agency that enforces consumer protection law, settled with Everalbum, the company behind the Ever photo app, after it applied facial recognition to users' photos without proper consent. The settlement didn't just require Everalbum to delete the improperly collected photos, it required the company to delete the facial recognition models and algorithms it had built using that data. It was the first time the FTC had gone after the model itself, not just the underlying dataset, setting a precedent immediately recognized as significant: intellectual property built on tainted data isn't safe from disgorgement (FTC, 2021).


In 2022, the FTC went further with WW International (formerly Weight Watchers) and its subsidiary Kurbo, which had collected health data from children as young as eight without proper parental consent. The settlement ordered the companies to delete the improperly collected data, destroy any algorithms or models built from it, and pay a $1.5 million penalty (FTC, 2022).


This is often described as algorithmic disgorgement: a legal penalty requiring companies to delete unauthorized data and any artificial intelligence models or algorithms trained on that data. The pattern in both cases makes the concept concrete: the model inherits the sin of its data. If the consent underneath a dataset doesn't hold up, everything built on top of it, the embeddings, the fine-tuned model, the product shipped to customers, becomes exposed to the same claim.


India's own law is catching up, and buyers shouldn't wait for it

India's Digital Personal Data Protection Act, passed in 2023, gives individuals ("Data Principals") the right to withdraw consent at any time, and obligates the company holding the data ("Data Fiduciary") to erase personal data once consent is withdrawn or the stated purpose is fulfilled (Carnegie Endowment, 2023). Europe's GDPR carries similar erasure and purpose-limitation obligations, and has for years.


Here's the detail worth knowing if you're building on Indian data: the DPDP Act's substantive compliance machinery isn't fully live yet. India's Data Protection Board became operational in November 2025, but the core obligations, consent notices, erasure rights, breach reporting, don't come into force until May 2027 (Matters.ai, 2025). That gap is not a reason to relax. It's a reason to move early. How the DPDP Act will treat data collected before its obligations come into force is still untested, and there's no case law yet to say how retrospectively it will be applied. But the FTC's disgorgement cases already show what regulators elsewhere are willing to do to downstream models, and India's own enforcement posture is likely to harden, not soften, over time. Waiting for enforcement to force good practice is a bet against your own downstream product.Training data collected carelessly is a loan that can be called

That's the frame worth sitting with. A dataset with weak consent isn't just a legal liability sitting in a folder somewhere, it's embedded in whatever gets built from it. If a regulator, a court, or simply a worker exercising their right to withdraw consent forces a repossession, what gets clawed back isn't just the raw footage. It can be the model, the months of training time invested in it, and the products already deployed on top of it.


If you're buying robotics data, you should care about chain of custody, consent, purpose limitation, withdrawal mechanisms, auditability and provenance (the documented history of where data came from, how it was collected and processed, and under what permissions), not merely volume and annotation quality. A buyer who can't answer "was this person told what this footage becomes, and can they get out?" is buying exposure along with the dataset.In practice, that means asking specific questions before signing off on a dataset: Was consent captured on record, in the worker's own language, rather than a signed form they couldn't read? Does the data come with a clear, time-stamped chain from the individual to the vendor to you? Can the vendor show what happens if a worker withdraws consent after the data has already been used in training? Is there a documented purpose limitation, so the data isn't quietly reused for something the worker never agreed to? And is any of this auditable by a third party, or does it rest entirely on the vendor's word?


What good practice actually looks like


At DesiCrew Solutions, we've collected and labelled data for global technology companies since 2007, across text, voice, image and video, across Indian districts and outside Indian geography too. The honest lesson from nearly two decades of this work is that sourcing data was never the hard part. What's hard, and what's worth doing properly, is explaining consent to a worker in their own language, building a process that can withstand an audit, and giving people a real, working mechanism to say no or to withdraw later. It's slower. It costs more upfront. But it's the only version of this business a buyer never has to apologise for.

For physical AI, provenance and quality aren't separate concerns. They reinforce each other.


That's the bridge we're trying to build: between technology companies and the real-world environments where AI is trained, tested and deployed.


Consent, in this industry, isn't a compliance checkbox. It's a data-quality and asset-quality issue, and increasingly, it's the difference between owning a model and merely renting one until someone calls in the loan.


__


(Saloni Malhotra is the Founder of DesiCrew Solutions. More on www.desicrew.in.)


Comments


bottom of page