Residential Proxies for Market Research, Explained
Open the same retailer’s product page from Warsaw and from Toronto. You may not be looking at the same page. The currency differs. So does the assortment, the promotion banner, the delivery estimate, and sometimes the price itself.
Most research teams never see the second version.
Google formalised this pattern in October 2017. It stopped tying results to the domain someone typed and began serving the country service that matched where the person actually was. Google’s announcement gave the reason: about one search in five relates to location. Type google.co.uk from Madrid after that change and Spanish results come back.
A team collecting from one office therefore records one localised rendering of every site it studies. The dashboard still looks clean. Nothing on it flags that the numbers describe London and nowhere else.
What the price-personalisation research found
Perhaps those regional differences are trivial. Two bodies of evidence say otherwise.
Aniko Hannak and four colleagues at Northeastern University built a method for separating real personalisation from measurement noise. They ran it across 16 major e-commerce sites, using the accounts and cookies of more than 300 real users. Their 2014 paper for the Internet Measurement Conference found evidence of personalisation on nine of the sixteen. Some sites reordered which products appeared. Others changed the prices attached to them.
The practice has spread since. In January 2025, FTC staff released the first results of a surveillance pricing study built on 6(b) orders to eight intermediary firms. Staff reported that inputs as fine-grained as a shopper’s precise location, browsing history, mouse movements and abandoned cart contents can feed the price a retailer displays.
Precise location leads that list.
A researcher who samples one location is not being cautious. They are measuring one stratum and reporting it as the whole.
Where a vantage point comes from
This is why the network layer stopped being plumbing and became a design decision.
A datacenter address belongs to a hosting provider. Those providers publish their ranges, and detection systems recognise them easily, so plenty of sites treat that traffic differently — generic pages, throttling, or an outright block. For a study whose entire point is the local experience, a generic page does more damage than a missing observation. It looks like data.
A residential address is one an internet service provider handed to a household, so a request carrying it arrives as ordinary local traffic. Networks of residential proxies route requests through ISP-assigned household addresses across dozens of countries. Nothing about the arrangement is a disguise. The connection really does originate where it claims to.
Carrier-assigned mobile addresses form a third tier, and they cost accordingly. They earn their premium on the narrow set of tasks that touch mobile-specific experiences, and almost nowhere else. Matching the tier to the task is a cost decision and a data-quality decision at the same time.
The rules already exist, and they are not the ones people quote
robots.txt became an IETF Proposed Standard in September 2022. RFC 9309 sets out the syntax, tells crawlers not to use a cached copy for more than 24 hours, and requires parsers to handle at least 500 kibibytes. One line in it deserves more attention than it gets: “These rules are not a form of access authorization.” Honouring robots.txt is courtesy and good practice rather than a legal shield, and ignoring it is not a crime either.
In US law, the Ninth Circuit took up the Computer Fraud and Abuse Act in hiQ Labs v. LinkedIn. On remand after the Supreme Court’s Van Buren decision, the court ruled in April 2022 that scraping publicly available pages — pages that need no account — probably does not amount to access “without authorization.” That reasoning covers the CFAA and nothing else. Contract terms, copyright and trespass claims run on separate tracks.
Europe draws its line around personal data instead. The European Data Protection Board adopted Guidelines 03/2026 on web scraping in July 2026. The central point travels well beyond the generative AI context they address: publishing something online does not release it for any later purpose. Consent rarely works as a basis at scale, and legitimate interest demands a documented three-part assessment.
Three sources, one practical rule. Stay on public pages. Collect prices, stock, listings and public content. Leave personal data alone.
Three checks that decide whether the data is any good
The ICC/Esomar International Code, which Esomar revised in July 2025 and which more than 60 associations across 50-plus countries recognise, is blunt about method disclosure. Article 7(b) tells researchers to state their limitations openly, including “potential gaps in data sources or population representation.” Article 9(b) goes further and requires enough technical documentation for an outsider to validate published findings.
Three measures make that promise keepable.
Coverage. What share of targets, in the markets that matter, produced a usable observation in the latest run? A network strong in five countries cannot carry a thirty-country study, and the shortfall turns into a silent blind spot.
Freshness. How old is the newest data point at the moment someone makes the decision? Studies age in drawers. People then quote the stale numbers in meetings six months later as though nothing moved.
Agreement. How often does a sampled observation match what a person in that market sees on the same page? This is the check that catches a pipeline treating failed requests as genuine empty results, which quietly corrupts every average downstream.
Where this goes next
A difference between markets is not by itself evidence of wrongdoing, and researchers should resist writing it up as one. Regulation (EU) 2018/302 stops traders from blocking access or applying different general conditions because of a customer’s nationality, residence or place of establishment. It does not force identical prices across national storefronts. A price gap between two EU country sites is a fact to record, not an accusation to make.
Meanwhile the personalisation keeps deepening. More sites now assemble what they show using machine learning, which widens the distance between what any single observer sees and what actually exists across a market.
Both frameworks above remain unsettled. The FTC study is still running, and the EDPB only finalised its scraping guidelines in July 2026. Anyone building a collection programme now is building it on ground that will shift.
There is no longer one internet to study. There are many overlapping local ones, and every site assembles its own version on request. Teams that name the network layer out loud can design around it. Teams that leave it invisible will eventually stand in a meeting defending a number nobody can reproduce.
Observer Voice is the one stop site for National, International news, Sports, Editor’s Choice, Art/culture contents, Quotes and much more. We also cover historical contents. Historical contents includes World History, Indian History, and what happened today. The website also covers Entertainment across the India and World.