Web scraping is using software to read web pages and pull out structured data automatically.
Key points
- A scraper fetches a page over HTTP, parses the HTML and extracts specific elements, often on a schedule [1].
- The Robots Exclusion Protocol, standardized as RFC 9309 in 2022, lets site owners say which paths automated clients should not crawl [2].
- An official REST API is usually more stable and clearly permitted than scraping the same data from HTML.
- Scraping Personal Data does not make it free to use; under GDPR a lawful basis and transparency duties still apply, even for Publicly Available Data [3][4].
- Responsible scrapers respect robots.txt, site terms and Rate Limiting, and identify themselves honestly.
How web scraping works
A scraper does what a browser does, minus the person. It sends an HTTP request for a page, receives the HTML, and uses rules to find the parts it wants, such as a product price, a table row or a company description [1]. The extracted fields are written to a file or database, and the process repeats across many pages. Simple scrapers read static HTML; others drive a headless browser to render pages built with JavaScript. Scraping differs from crawling, which follows links to discover pages, though the two are often combined. Common legitimate uses include price monitoring, research, archiving, search indexing and pulling public details about a company from its own website to support research tasks such as understanding its Firmographics or Technographics.
Rules and etiquette
Site owners can publish a robots.txt file stating which parts of a site automated clients may visit. RFC 9309 standardized this Robots Exclusion Protocol in 2022, while noting that it is a request, not an access control [2]. Well-behaved scrapers honor it, limit their request rate so they do not burden the server, and identify themselves with a clear user agent. Site terms of service may restrict automated access further, and courts in different countries have treated such terms differently, so legal advice is worth getting for any large project. Where a site offers an official REST API, using it is usually the better choice: it is documented, stable, and comes with explicit terms, API keys and published Rate Limiting.
Personal data and privacy
The hardest questions arise when scraped data identifies people. Under GDPR, information about an identifiable person is Personal Data whether or not it is public, and collecting it requires a lawful basis such as Legitimate Interest [3]. When data is not collected from the person directly, Article 14 requires telling them who holds it, why and where it came from, within a reasonable period and at the latest within one month [4]. Data Minimization also applies: collect only the fields you need, and delete them when they are no longer needed. For sales teams, the practical takeaway is that any data about individuals, however it was gathered, belongs under the same privacy controls as the rest of the CRM (Customer Relationship Management), including honoring objections and the Right to Erasure.
Related terms
Outreach without the busywork.
PineLead finds new B2B prospects every day, qualifies them against your criteria and writes the first email in your voice. You approve — PineLead sends.
Start free with 100 credits →