Tips for building a crawler from scratch
Notes from the team behind UnGovrBot, written for r/webscraping and for anyone building a scraper with a coding assistant this weekend. The architecture is the easy half. The half that sinks projects is the law, and it is the half almost nobody designs for.
UnGovr crawls government websites for public records: agendas, minutes, budgets, codes. Our crawler is called UnGovrBot, and what it promises site owners is at https://www.ungovr.org/crawler. This page is for the other audience, people building one.
We keep this at the level of design. It names no tools and gives no recipe for getting past anything. What it does give is the order we would build things in if we started again, and the question we wish we had asked on day one: who answers for what this program does?
1. The architecture, in nine parts
A crawler that lasts is nine small parts with clear seams, not one loop. Most trouble comes from two of them being fused: fetching with extraction, or policy with retry logic.
| Part | What it does | The mistake to avoid |
|---|---|---|
| A reason not to crawl | Looks for an API, a bulk file, a feed, a sitemap or an open-data portal before any page is fetched. | Scraping HTML from a publisher that offers the same data as a download with a licence attached. |
| Frontier | Holds what to fetch next as one queue per host, each with its own budget. | One global list, which lets a single large site starve the rest or take all the load. |
| Scheduler | Owns politeness: one or two connections per host, a delay between requests, back-off on a 429 or 503, quiet hours in the host's own time zone. | Putting the delay inside the fetcher, where a retry quietly ignores it. |
| Identity | Says who is asking: a named User-Agent with a URL, a contact address, reverse DNS, and signed requests where the site can verify them. | An anonymous script with no way to reach its operator. A site that cannot ask you to slow down can only block you. |
| Policy gate | One function, called before every fetch, that answers "may I, this way, and on what basis" from data. | Scattering the answer across settings, retries and comments. |
| Fetcher | Gets the bytes the cheapest way that works, with conditional requests so an unchanged page costs almost nothing. | Starting every request in a full browser. |
| Store | Keeps the raw response, its headers and the time, addressed by a hash of the content. | Keeping only what you extracted. The day your parser changes, you crawl everything again. |
| Extractor | Turns stored bytes into records, versioned, and able to run again over old bytes. | Parsing during the fetch, so one bad selector loses the page. |
| Decision log | Records why each fetch was allowed, slowed, changed or refused. | Logging errors only. The questions that matter later are about the fetches that worked. |
Two smaller things save more time than any library choice. Detect traps early (calendars that go on forever, session identifiers in links, filter pages that multiply), and treat a URL and its content hash as two different keys, because the same document will reach you at five addresses.
2. A fetch ladder, and where it has to stop
Most pages come back from a plain request. Some do not, and the reason matters more than the status code. A page that needs a browser to render has not refused you. A site that answers 429 is asking you to slow down. A challenge page from a filter in front of the site is a guess about what you are, often made by a default setting nobody at the site chose.
So a fetcher that lasts is a ladder, cheapest rung first.
| Rung | Use it when | What changes |
|---|---|---|
| 1. Plain request | Always first, under your declared identity. | Nothing. |
| 2. Slower | A 429, a 503, timeouts. | Your pace. Days, if that is what the host can bear. |
| 3. A real browser, same identity | The page is a script shell: a 200 with nothing in it. | How the page is rendered. Not who you say you are. |
| 4. Another door | Still nothing. | The route: an API, a bulk download, a feed, or an email to the webmaster. |
| 5. Anything that changes who you appear to be | This is the rung people ask about. | Your identity. From here on the decision is legal, not technical. |
Escalate against a non-answer. Stop at a refusal.
A script shell, a rate limit and a generic filter's guess are non-answers: nobody decided anything about you. A login, a CAPTCHA, a letter telling you to stop and a block aimed at you are refusals. A CAPTCHA is a question addressed to a person, and a program that answers it has stopped being a crawler in any sense a court will recognise.
Why would anyone build the fifth rung at all? Because default filters turn away named, signed, slow crawlers from documents the publisher is required to publish, and the site's own staff often do not know it is happening. That is a real problem, and section 4 walks through one. But whether you may climb is a question about a place, a barrier and a document. In the research behind LexLint, as of October 3, 2026, getting past a technical barrier is not coded as plainly permitted in any jurisdiction it covers.
For what it is worth, here is where our own crawler sits. It identifies itself, signs its requests under a published internet standard, and may retry a bot challenge through its own tiers, from a plain request up to a full browser. A CAPTCHA is where it stops, and there are sites it will only ever visit as the declared crawler. What one operator does on that rung is a risk it has chosen to carry in places where it has read the law. It is not a line the law has drawn, and your project needs its own answer.
3. Your crawler is you
No legal system we know of treats a program as a party. The law looks through the software to the person or company that runs it. United States federal law says so in terms: a contract formed by an "electronic agent" stands "so long as the action of any such electronic agent is legally attributable to the person to be bound" (15 U.S.C. § 7001(h)).
Courts say it more bluntly. When Air Canada argued that it was not responsible for what its website chatbot told a customer, a British Columbia tribunal in 2024 called that "a remarkable submission" and held the airline to the answer, because the chatbot "is still just a part of Air Canada's website".
This cuts both ways, and both matter to a scraper. Your rights travel with your agent: if you are entitled to read a document, you may send a program to read it. So does your liability: "the script did it" and "the assistant wrote it" are not defences, and the terms a script accepts, the letter it ignores and the barrier it gets past are yours.
The code you downloaded is yours too
The same rule reaches software you did not write. Running an open-source scraper whose defaults you never read is not much of a defence. A contract claim, a copyright claim and a data-protection duty do not turn on what you knew about the library, and where intent does matter, you are the one explaining the settings you shipped.
Nor can you hand the bill back to the authors. The common open-source licences all say so, usually in capital letters. The MIT licence: "IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY". The Apache licence: "You are solely responsible for determining the appropriateness of using or redistributing the Work". The GNU General Public License: "THE ENTIRE RISK AS TO THE QUALITY AND PERFORMANCE OF THE PROGRAM IS WITH YOU". Whoever was harmed will look for the person who ran the program, and the licence has already told you the authors are not standing behind you.
So do not assume a project is lawful to run because it is popular, or unlawful because it looks aggressive. Read what it does by default: whether it reads robots.txt, how it identifies itself, what it does when it meets a challenge, and what it stores. Then check that behaviour against the places you run it, which is the job section 7 describes.
4. When the law is on the crawler's side
Here is a case every crawler of public records meets sooner or later. California's open-meetings law, the Brown Act, requires a local body to post its agenda at least 72 hours before a regular meeting. A city, county, special district or school district with a website must post it there, in an open format that meets three tests the statute spells out: "Retrievable, downloadable, indexable, and electronically searchable by commonly used internet search applications"; "Platform independent and machine readable"; and "Available to the public free of charge and without any restriction that would impede the reuse or redistribution of the agenda" (Government Code § 54954.2).
Now the city moves its website behind a content delivery network and leaves the default bot filter switched on. The agenda is still there for a person with a browser. A named, signed, slow crawler asking for the same file gets a challenge page. The meeting is in three days.
Three different things can stand between a crawler and that agenda, and the law does not treat them alike.
| What you meet | What it is | Where the law stands |
|---|---|---|
| A robots.txt rule | A published request. | The standard itself says its rules "are not a form of access authorization" (RFC 9309). A New York federal court held in December 2025 that such a file controls access no more than a "keep off the grass" sign controls access to a lawn. In California no statute or reported decision gives it legal force, which means untested, not safe. In the European Union a machine-readable reservation can remove the text-and-data-mining exception, so there it carries real legal weight. |
| A default bot filter | A third party's guess about what you are. | Unlit. In July 2026 the same New York court let anti-circumvention claims go forward over a challenge system that had to be passed before pages were served. That ruling concerned copyrighted work and decided only a motion to dismiss. Our own research codes getting past a barrier in California as restricted, not permitted. |
| A CAPTCHA, a login, a letter, a block aimed at you | A decision, addressed to you. | Stop. Carrying on after being told personally to stop is the fact pattern behind real liability. |
Our reading of the agenda case is this. The statute puts the duty on the agency, and the public, with whatever tools it uses to read, is who that duty runs to. A request file cannot repeal a statute. A filter the city never configured is not the city refusing anyone. So reading the posted agenda with a program is, in our view, what the statute contemplates.
That is a reading, not a ruling. No court has decided a Brown Act agenda against a bot filter, and the statute tells the agency what to publish. It does not say what a member of the public may do to a filter that is in the way. So the order of moves matters more than the argument:
- Be identifiable. A named, signed request is the difference between a reader and an intruder when someone looks at the logs later.
- Write to the clerk and the webmaster first. Cite the section. Ask for an allow-list entry or a feed. The block is usually a default, and a default can be changed by the people who own the site.
- Take only what the statute makes public, at a pace a person could match.
- Write down the basis before the fetch, not after the letter.
- Stop the moment anyone at the agency says stop. That turns a non-answer into a refusal, and the analysis in the table changes rows.
None of this travels. It is a California statute about one kind of document. The same crawler, pointed at a company's site or at a government site in another country, starts again from nothing.
5. Three times it looked just as clear, and was not
The agenda case feels obvious: public document, no login, a law that says publish it. The cases below felt the same way to the people in them. None of them is us, and none of them is a household name.
A startup scrapes public profiles, and wins in the appeals court
- What they did
- hiQ Labs, a data analytics company, scraped LinkedIn profiles that anyone could see without logging in.
- Why it looked fine
- In April 2022 the Ninth Circuit held that scraping pages open to the public likely does not violate the federal Computer Fraud and Abuse Act. It is still the case people cite to say scraping public data is legal.
- How it ended
- In December 2022 the case ended with a $500,000 judgment against hiQ and a permanent injunction barring it from scraping LinkedIn, logged in or not, and from using or distributing the collection software it had built. The computer-crime question was never what decided it. LinkedIn's user agreement prohibited scraping, and hiQ had also used fake accounts to reach logged-in pages. The judgment was agreed between the parties, so it sets no precedent, but it is how the best-known win for scrapers finished.
A rental-listings startup collects public classified ads
- What they did
- RadPad, an apartment-rental listing company, scraped rental listings from Craigslist and emailed the people who had posted them.
- Why it looked fine
- The listings were public, free to read, and posted by people who wanted them seen.
- How it ended
- On April 13, 2017 a federal court in California entered a $60.5 million judgment: $40 million under the federal anti-spam law for 400,000 emails to addresses taken from the listings, $20.4 million for copyright in the scraped postings, and $160,000 for breaking the site's terms of use. RadPad was insolvent by then and its lawyer had withdrawn, so nobody argued its side. The scraping was the small number. What the data was used for was the large one.
One person notices that a public web address answers without a password
- What they did
- In June 2010 Andrew Auernheimer and a collaborator found that an AT&T web address returned an iPad owner's email address when given the device's identifier. A script guessed identifiers and collected about 114,000 addresses, which were passed to a reporter.
- Why it looked fine
- No password, no login, no barrier. The appeals court later noted that "no evidence was advanced at trial that the account slurper ever breached any password gate or other code-based barrier".
- How it ended
- A federal prosecution. A jury convicted him of conspiracy to violate the Computer Fraud and Abuse Act and of identity fraud, and the court sentenced him to 41 months in prison. On April 11, 2014 the Third Circuit threw the conviction out because the trial had been held in the wrong state. It did not decide whether the access was a crime. "The server answered" was not enough to keep a prosecutor, or a jury, from treating it as one.
Three different bodies of law, and none of them is the one a developer checks first. A contract. What you did with the data. A criminal statute read by a prosecutor rather than by an engineer.
6. Keep the rules as data
The policy gate from section 1 needs something to read. We suggest three groups of rules per jurisdiction, kept as plain data with a date on it, so that a change in the law is an update to a file and not a code review.
| Group | The question | Examples |
|---|---|---|
| Access | May I fetch this, this way? | What weight robots.txt carries. Whether terms bind you without a click. What a login changes. What follows a CAPTCHA, a block or a letter. |
| Content | May I keep it and use it like this? | Copyright exceptions and whether a site can opt out of mining. Database rights. Rules on quoting and on news snippets. Personal data, which stays personal when it is public. |
| Evidence | What must I be able to show? | What to log for each fetch, how long to keep it, and how fast you must report if your own store is breached. |
The second group is the one that surprises people. In December 2024 the French data protection authority fined KASPR, a company whose paid browser extension collected contact details from LinkedIn profiles, 240,000 euros. When people asked where their details had come from, the company told them "publicly accessible sources". Under European law that is not an answer: the authority found no lawful basis for details people had restricted, a retention period that ran too long, and notice that came years late and only in English.
Here is the same crawler in two places, from LexLint's records as of September 30, 2026. The values are its own vocabulary.
| Rule | California | European Union |
|---|---|---|
| Open page, no login | permitted | permitted |
| Page named in robots.txt | unsettled | restricted |
| Legal weight of robots.txt | non-binding notice | statutory |
| After getting past a technical barrier | restricted | restricted |
| After a letter telling you to stop | restricted | unsettled |
| Mining for commercial use | unsettled | allowed unless the site opts out |
| Personal data on public pages | minimize | minimize |
As a file your gate can read, one place looks like this. The access and content values are LexLint's for California. The evidence block is our suggestion, not a legal requirement.
And the gate that reads it stays small. The point of the sketch is the shape: every refusal has a reason, and the only way past a default is a decision a person made and wrote down.
7. Where the rules can come from
We had to answer these questions for our own crawler, in every place it runs. The answer became a research library of software law, and then a tool anyone can use: LexLint, which is in pilot and free to use. It works like a linter. You tell it what your project does and where, and it returns findings with citations and dates.
For a crawler, three touch points are enough.
- Before you build. The Scraping Checkup at https://lexlint.io/scraping/checkup asks up to twelve questions and needs no account. If you work with a coding assistant, paste
run https://lexlint.io/first-runinto it and it will read your project and run the lint with you. - In the project. Declare the activity
crawls_web(andtrains_modelsoraggregates_contentif they apply) and every place involved: where you are, where the crawler runs, and where the sites are. A place you leave off is not checked. - In the gate. Read one place's rules with the
get_lawtool, once, and cache the result with its date. It returns the access scenarios, the weight of robots.txt and the mining opt-out status used in the table above. Do not call it per page: the free allowance is 50 requests a day.
The same lint works on a project you downloaded. Run it in that project's directory before you run the code. Your assistant reads what the project does, you confirm it, and the findings are the law that applies to those activities in the places you name. It is not a line-by-line audit of someone else's code, so read the defaults yourself as well.
Two limits, stated plainly. LexLint reports exposure and never says a fetch is allowed, because a lint cannot know that. And it is research, not a lawyer: a real project with real risk takes its findings to one. The longer treatment of everything on this page is the Scraping Brief at https://lexlint.io/scraping, which is free to read.
8. What to log
When a letter arrives, or a regulator writes, the question is what you knew and when. A crawler that cannot show the robots.txt it read on the day is arguing from memory.
| Keep | Because |
|---|---|
| The request: address, time, status, content hash | It is the fact everything else hangs from. |
| The identity you presented | It shows you did not hide, which is most of good faith. |
| The robots.txt you read, as bytes, and the rule that matched | Files change. Yours is the copy that was true when you fetched. |
| The rung used, and why it changed | A retry is a decision. Record who or what made it. |
| The site's terms and licence as you saw them | Contract claims turn on what you were shown and when. |
| The place, the rule set and its date | It shows the basis existed before the fetch. |
| Every message asking you to stop, and when you did | From that moment the analysis changes, and the clock is yours. |
Then give the log a retention period and hold to it. A fetch log is full of addresses and, sooner or later, people's names, which makes it personal data in much of the world. If your own store is breached, reporting deadlines start, and they differ by place: https://lexlint.io/clock draws them on one axis.
The short version
- Look for a door before you build a ladder.
- Say who you are, and sign it.
- Read what you download. Its defaults are your conduct, and its licence says the risk is yours.
- Escalate against a non-answer. Stop at a refusal.
- The email comes before the workaround.
- Rules are data with a date. Decisions are made by people and written down.
- The log is evidence. Keep what you would want to show, and no longer than you need it.
Questions, corrections, or a site of yours we should be reading differently: crawler@ungovr.org.
Sources
- California Government Code § 54954.2, agenda posting: https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=GOV§ionNum=54954.2
- RFC 9309, Robots Exclusion Protocol: https://www.rfc-editor.org/rfc/rfc9309.html
- 15 U.S.C. § 7001(h), electronic agents: https://www.govinfo.gov/content/pkg/USCODE-2023-title15/html/USCODE-2023-title15-chap96-subchapI-sec7001.htm
- Moffatt v. Air Canada, 2024 BCCRT 149: https://decisions.civilresolutionbc.ca/crt/crtd/en/525448/1/document.do
- The MIT License: https://opensource.org/license/mit
- Apache License 2.0, sections 7 and 8: https://www.apache.org/licenses/LICENSE-2.0.txt
- GNU General Public License 3.0, sections 15 and 16: https://www.gnu.org/licenses/gpl-3.0.txt
- The two New York rulings on robots.txt and challenge systems (Ziff Davis v. OpenAI, December 15, 2025, and Reddit v. Perplexity, July 31, 2026), set out with their sources: https://lexlint.io/scraping/defensible-barriers
- hiQ Labs v. LinkedIn, Ninth Circuit, April 18, 2022: https://cdn.ca9.uscourts.gov/datastore/opinions/2022/04/18/17-16783.pdf
- The hiQ consent judgment, December 2022, as reported by Proskauer: https://newmedialaw.proskauer.com/2022/12/08/hiq-and-linkedin-reach-proposed-settlement-in-landmark-scraping-case/
- Craigslist v. RadPad, No. 16-01856 (N.D. Cal. April 13, 2017), as reported by Proskauer: https://newmedialaw.proskauer.com/2017/04/17/craigslist-garners-60-million-judgment-against-radpad-in-scraping-dispute/
- United States v. Auernheimer, No. 13-1816 (3d Cir. April 11, 2014): https://www2.ca3.uscourts.gov/opinarch/131816p.pdf
- The French data protection authority's decision on KASPR, December 5, 2024: https://www.cnil.fr/en/data-scraping-kaspr-fined-eu240000