Why JD.com Is Giving Away Free Helmets
In India, they paid $2.60 an hour for this exact footage.

Opening
Reader, on August 3rd, JD.com unveiled a smart helmet for its delivery riders. It takes voice orders, analyzes veteran riders’ actual driving records to guide routes, and sends a rescue signal if the rider falls. JD.com says it will hand these out free, on a priority basis, to 150,000 full-time riders.
Most articles framed this around “safety” and “efficiency.” But something else caught my eye: this helmet has a camera that reads the surrounding environment.
Last issue, I told you about the textile workers in Karur, India — the ones who strapped GoPros to their foreheads to film themselves folding clothes, earning roughly $2.60 an hour. Here’s the punchline: China just found a way to stop paying that $2.60.
In the First Week of August, Three Companies’ Helmets Lined Up Side by Side
Here’s the feature list JD.com released for its helmet. Riders can process new orders by voice, take customer calls, and have requests like “leave it at the door” read aloud automatically — no need to tap a screen with wet hands. Instead of a map app, the helmet learns the routes experienced riders have actually traveled and gives turn-by-turn directions based on that. It even automatically checks hygiene conditions when riders pick up food at a restaurant. JD.com says it expects the helmet to cut route-finding time by roughly 3 minutes per delivery.
JD.com wasn’t first, though. Meituan had already distributed more than 1 million smart helmets by June 2025. Its version includes voice calls, fall-detection sensors, and automatic emergency response. Alibaba’s delivery arm (Ele.me at the time, now known as Taobao Shangou) had already rolled out a similar helmet as far back as late 2022.
China has more than 10 million delivery riders as of early 2025, according to figures reported by Xinhua. Which means three platforms have been pushing the same type of equipment into this market over the span of three years.
One thing worth flagging here: the safety benefits these helmets provide are real. Handling things by voice is genuinely safer than fumbling with a wet smartphone in a downpour. Fall detection and one-touch rescue requests are features that can change survival outcomes when accidents happen. This isn’t marketing rhetoric — it’s physically true.
But the three companies’ helmets share one more thing in common. Cameras, microphones, and location sensors are all fixed at the rider’s head height, moving through the city all day long. And all of them are free.
Why the Head, of All Places
What’s the hardest resource to get your hands on in robotics right now? It’s not models, and it’s not chips. It’s first-person footage shot at human eye level1.
Last June, researchers from Peking University, the National University of Singapore, MIT, UC Santa Barbara, and NVIDIA jointly published a paper called “HumanScale” that tackles this head-on. The experimental design is clean. They pre-trained2 robots on the same 5,000 hours, filled two different ways: one set was first-person footage shot by humans wearing head-mounted cameras, the other was actual driving and manipulation logs collected by robots.
The results ran against conventional wisdom. When given tasks they’d never seen during training3, the model trained on human first-person footage had roughly 20% lower loss. The gap was even more dramatic in physical robot tests. A robot with no pre-training scored 0% success on an unseen task, while a robot pre-trained on first-person footage scored 90%.
Why would that be? The paper’s explanation is simple. Robot data is made in labs, following a fixed script, only within reach of a stationary arm. Human footage, by contrast, moves through homes, streets, and workplaces; its motion is fluid, with little idle time. Even filming the same 100 hours, the number of distinct motion trajectories extracted from human footage was more than 5 times that from robot data.
So how scarce is this data, exactly? Ego4D, the standard first-person dataset used in academia, consists of 3,670 hours shot by 923 people across 9 countries — a scale gathered over years as a research project.
Now let’s set JD.com’s numbers next to that. Even under a very conservative assumption — 150,000 full-time riders filming just 1 hour each per day — that’s 150,000 hours a day. That’s 40 Ego4D’s worth of footage accumulating daily. It’s enough to fill the 5,000 hours HumanScale used for pre-training 30 times over, with room to spare. Scale that up to China’s entire rider population of 10 million, and the math moves into territory where comparison stops being meaningful.
Let’s convert this using Karur, India’s wage rate. At $2.6 an hour, that’s $390,000 a day, and over $140 million a year — based on 150,000 riders. That’s the scale of data-collection labor cost that never even made it onto anyone’s books to begin with.
Of course, in fairness, two things need to be flagged.
First, the paper’s limitations. As the authors themselves note, the comparison is capped at 5,000 hours (because publicly available real-robot data itself is scarce). Physical robot validation was done on only 3 tasks and 1 robot platform. Whether the advantage holds at scales of hundreds of thousands of hours remains unconfirmed. That said, performance kept improving log-linearly from 100 hours all the way to 5,000 hours (R²=0.86–0.94), which reads as a signal that “more input keeps yielding more gains.”
Second, the nature of the data differs. What HumanScale dealt with was mainly hand-based manipulation tasks. What a helmet camera captures is closer to urban navigation — the layout of alleyways, shortcuts between apartment buildings, the route to find and enter a storefront, waiting in front of an elevator. That’s raw material that maps more directly onto autonomous delivery vehicles and delivery robots than onto robotic arms. And, as it happens, autonomous delivery vehicles are exactly what’s spreading fastest in China right now.
The Line Item Missing From the Contract
This is where the difference between last issue and this one becomes visible.
The model in Karur, India had a name. There was a data-collection contract, an hourly wage, and the worker knew they were doing “AI training video work” right now. Exploitative, maybe, but visible. Because it was a transaction, it could be negotiated — and there was room for the rate to rise.
The model in China has no name. What the rider signed was a delivery contract; what’s on their head is safety gear. Data collection isn’t a separate side gig — it’s been folded into work they were already doing. No payment, no line item, and so there’s nothing to negotiate in the first place.
If last issue’s question was “who owns which layer,” this issue’s question is: what happens when that layer isn’t visible at all?
The evidence, actually, is already sitting in the product spec sheet. One of the JD.com helmet’s core features is “veteran-rider route guidance” — it analyzes the actual delivery records of top-performing riders and voice-guides new riders along the optimal route. Look closely, and this is a device that extracts a skilled worker’s experience into a model and redistributes it to unskilled workers. Nowhere does it say the veteran was paid separately for that know-how. This isn’t inference — it’s the company’s own stated feature description.
The timing isn’t coincidental either. China’s State Administration for Market Regulation released a draft amendment to the e-commerce law on July 4, strengthening platform responsibility for the safety, wages, and basic social security of outsourced workers like delivery riders, with a comment period running through August 4. China has roughly 300 million flexible-employment workers.
But the riders’ reactions are telling. A delivery worker named Zhang Liang, quoted by SCMP, works 12-hour days and earns 300 yuan (about $44). What he said he wanted wasn’t social insurance — it was “higher per-order fees and reasonable delivery windows.” If insurance premiums rise, less cash ends up in his pocket.
For platforms, this legislation reads as a signal that labor costs are going up. And the opposite curve is falling at the same time. China’s autonomous delivery vehicles topped 15,000 units deployed nationwide as of July, with operating approval in over 200 cities. A vehicle that cost more than 1 million yuan a few years ago has now dropped below 20,000 yuan. Hardware costs have fallen from $30,000 in 2021 to under $10,000 now.
JD.com chairman Liu Qiangdong, speaking at the APEC China CEO Forum in Beijing on June 21, said robots would take over deliveries and that 700,000 delivery riders could ultimately be replaced. He added that the company is preparing job-transition training in partnership with more than 120 schools nationwide.
There’s one more layer. In China, surveying roads and terrain and collecting geographic data isn’t something just anyone can do — it’s licensed, and geographic data collected by vehicles in particular is subject to separate controls. In effect, this closes off the path for foreign companies to scan Chinese streets and accumulate this kind of data. Which means the urban scan data generated daily above riders’ heads is, structurally, an asset only domestic Chinese platforms can accumulate. The question from last issue — who owns which layer — reappears here, but at the scale of the nation-state.
Let me finally draw an honest line. JD.com has only stated that it will use the data accumulated on the helmet to improve dispatch efficiency and safety management; it has never announced that the camera footage is used to train robots or autonomous vehicles. There is no public evidence yet directly connecting the two facts. What I’m pointing to isn’t causation but structural proximity. Sitting inside the same company, side by side, are: a device producing the world’s most valuable first-person raw video footage, a robotics business that needs exactly that raw material, and a CEO’s forecast that 700,000 people will no longer be needed.
Oswarld's Lens
I think the heart of this issue isn’t surveillance — it’s accounting.
There’s a pattern I’ve seen again and again while building GTM strategies. The most reliable way to eliminate a cost isn’t to cut it — it’s to erase the fact that it was ever a line item in the first place. In the India model, data collection was a line on the income statement. A line means someone can look at it, scrutinize it, and demand that it be raised. In the China model, data collection gets booked under the safety-equipment budget. Nobody calls it a data purchase.
Labor issues tend to turn dangerous not when compensation is low, but when a transaction isn’t recognized as a transaction. A low unit price at least gives you a starting point to demand a raise. An unnamed line item doesn’t even give you the words to make the demand.
And this isn’t just a story about someone else’s delivery market. Right here at home, work apps, in-vehicle terminals, wearable scanners in logistics centers, and recorded customer-service calls generate enormous volumes of work data every single day. Yet whether a company can repurpose that data for AI training, how far that repurposing can go, and how compensation for it should be calculated — these questions are rarely spelled out in employment contracts or subcontracting agreements.
When I help companies review new technology adoptions these days, there’s a question I’ve added to my checklist: “Does the contract include a clause on secondary use of the data this tool collects?” Most of the time, the answer is no. And “no” doesn’t mean it’s prohibited — it means nobody ever asked the question.
I run a newsletter that analyzes technology, economics, and the humanities every week, crossing between them. Digging into questions like this one is exactly what this newsletter is for.
Closing
Here’s the summary.
First, first-person video shot from human head height is right now the most expensive raw material in robotics. In the “HumanScale” experiment, this footage produced better generalization performance than an equal amount of real robot data.
Second, India bought that footage by paying an hourly wage for it, while China is folding it into safety equipment. The difference isn’t the unit price — it’s whether the transaction is visible or not.
Third, this equipment emerged at the exact crossing point where labor costs are rising and machine costs are falling. Both things are true at once: the helmet genuinely does improve safety, and the prospect of rider replacement has been stated publicly.
Last issue, I asked whether the Karur workers were being liberated or exploited. This issue’s question is a little different. What should we call a transaction that has no name, in exchange for no compensation?
💬 Among the work tools you use right now — company apps, in-vehicle terminals, call recordings, wearable scanners — is there one where you thought, “the company could probably use this as training data”? Tell me in the comments which tool it was and why you felt that way. If enough cases come in, I’ll pull them together in the next issue from the angle of Korean contract practice.
💬 Tell me in the comments about a tool where you thought, “this is probably being used as training data.” I’ll work it into the next issue. 📨 If you know someone who’s also thinking hard about data, labor, and contracts, please pass this piece along.
References & Further Reading
Primary sources
- Ma, J., Bi, J., et al., “HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining”, arXiv:2606.20521, 2026. Link ··· This is the backbone of today’s piece. Start with the table comparing the diversity of human video versus robot data item by item, and the passage on real-robot success rates of 0% versus 90% — it’ll immediately click why first-person video is so valuable.
- “HumanNet: Scaling Human-centric Video Learning to One Million Hours”, arXiv:2605.06747, 2026. Link ··· This is the original repository HumanScale drew its training data from. You get a real sense of scale: over 800,000 of the 1,000,000 hours are egocentric video.
- “China’s tech giants race to put AI on delivery riders’ heads”, South China Morning Post, 2026. Link ··· An article that compares the helmets of JD.com, Meituan, and Alibaba side by side. It’s the starting point of today’s piece.
- “Stronger social security? Proposed law sparks a pushback by gig workers”, South China Morning Post, 2026. Link ··· The key passage is where riders themselves say they’d rather have per-delivery fees than social insurance. It shows the paradox of protective legislation.
- “China’s Delivery Revolution”, The Wire China, October 2025. Link ··· Tracks in numbers how the price of autonomous delivery vehicles fell from 1,000,000 yuan to 20,000 yuan. Worth reading alongside the labor-cost curve.
Background
- Grauman, K., et al., “Ego4D: Around the World in 3,000 Hours of Egocentric Video”, CVPR, 2022. Link ··· The de facto standard dataset for first-person video. Knowing the scale — 923 people, 3,670 hours — makes the arithmetic in this issue much clearer.
- “‘700,000 delivery riders won’t be needed anymore’… Warning from JD.com’s chief in China”, Seoul Economic Daily (Sedaily), June 2026. Link ··· Lays out Richard Liu’s remarks at the APEC forum alongside the context of China’s 320,000,000 gig workers.
- Restrictions on geographic data in China, Wikipedia. Link ··· A quick overview of China’s surveying and mapping data regulations. Helps explain why this data becomes an asset exclusive to domestic platforms.
Related past issues worth reading together
- GoPros on Foreheads: The People Raising Robots ··· This is the prequel to today’s piece. Here you can find where the $2.6 hourly wage in Karur, India comes from, and how the framing of “layer ownership” was constructed.
- What the Luddites Smashed Wasn’t the Machine ··· The structure in which workers end up contributing to training the very system that will replace them existed 200 years ago too. Read it alongside today’s piece and the pattern becomes visible.
📝 Glossary
Footnotes
-
Egocentric video: Video shot from a camera mounted at head or eye level. Unlike third-person footage, it captures the exact moment hands touch objects and the actual angle a person sees from — making it far more useful for teaching robots movements. ↩
-
Pretraining: The stage of building a model’s basic competence using large amounts of general data before teaching it a specific task. It’s similar to having someone learn how to hold a knife and handle fire before teaching them to cook. ↩
-
Out-of-distribution generalization: The ability to perform properly in situations never seen during training. This capability is decisive for a robot to be useful in the real world beyond the lab. ↩




Your take shapes the next issue
What resonated most in this issue, or where has your experience been different?