Copart scraper or API: what scraping Copart really costs
What a Copart scraper really needs: proxies, layout fixes, images, monitoring and a sold archive you cannot backfill. The cost structure, and when to build.

A Copart scraper is cheap to write and expensive to keep. A developer with a headless browser can pull lot pages into JSON in a day or two; the cost arrives afterwards, as proxies, bot protection, layout changes, image handling, monitoring and the hours spent fixing it, every month, for as long as your product depends on it. And the one thing money cannot buy later, an archive of sold lots with final prices, only exists from the day your scraper started.
We have collected Copart and IAAI data every day for years, so this is written from the maintenance side. It is not a how-to for getting around anyone's defences: we do not describe that here. It is a cost breakdown, with the assumptions written out so you can plug in your own numbers, and an honest list of cases where building your own collector is the right decision.
Key takeaways
- The first version of a Copart scraper is the cheap part; maintenance is a recurring cost with no end date.
- Bot protection, layout changes and image URLs are the usual breaking points, and you notice them only if you monitor for them.
- Sold prices have to be captured when each auction closes; a scraper that starts today has no history before today, and every outage leaves a permanent gap.
- Refreshing a large catalog by re-reading pages costs orders of magnitude more requests than reading a change feed.
- Building your own is reasonable for narrow, short-lived or unusual needs, and for sites no provider covers.
What a Copart scraper has to do
"Scrape Copart" sounds like one job. In production it is six, and each one fails in its own way.
- 1DiscoverFind every lot: search pages, sale lists, new arrivals. Miss a page and those lots never exist for you.
- 2FetchLoad each lot page or the data behind it, within rate limits, through proxies, with retries.
- 3ParseTurn labels into fields: VIN, odometer and its status, damage, title, condition, bids, dates, location.
- 4NormaliseMap wording to stable values, convert times to UTC, split vehicle from lot, dedupe relists.
- 5Track changesRe-read active lots for bids and Buy Now, notice when a lot closes, capture the final price.
- 6OperateStore photos or links, monitor every step, alert, fix, redeploy. Forever.
Most teams estimate the first three steps and discover the other three in production. Normalisation alone is a project if you want IAAI next to Copart: different labels, different identifiers, different page structure (the Copart vs IAAI comparison lists them).
Bot protection, proxies and terms of use
Large auction sites protect themselves against automated traffic, and Copart and IAAI are no exception. In practice a scraper at catalog scale meets rate limits, blocked IP ranges and browser challenges, and the rules change without notice. Teams respond with proxy pools and browser automation, which turns into a running cost and an arms race whose pace someone else sets. We are not going to describe how to defeat any of those measures, and we would not build a business on doing so.
There is also the contract question. Both auctions' terms of use restrict automated collection from their sites, and the legal position of scraping differs by country and depends on what you collect and how you use it. That is a conversation for your lawyer, not for an engineering blog, but it belongs in the cost column: a plan whose central dependency can be withdrawn at any time is a risk you are pricing in, whether you write it down or not.
Where Copart scrapers break
Here are the incidents we plan for, roughly in the order a new scraper meets them. None of them is exotic; together they are why "it worked last week" is the most common sentence in scraper maintenance.
| Incident | What you see | How you notice | Typical fix |
|---|---|---|---|
| Page layout or data structure changes | Empty fields, wrong values in the right fields | Only with field-level validation | Rewrite selectors, re-parse affected lots |
| IP range blocked or challenged | Error pages, timeouts, sudden zero results | Success-rate alerts | Infrastructure work, more proxy spend |
| Site maintenance at night | Gaps in a run | Gap detection per hour | Re-run the window, if the data is still there |
| New price format or currency display | Parsing errors or silently wrong numbers | Range checks on prices | Parser fix plus data repair |
| New field appears (e.g. a new badge) | Nothing: you just do not have it | Someone compares with the site by hand | Schema change, backfill impossible |
| Image URLs change or expire | Broken photos on your site | Image 404 rate | Re-fetch or re-host media |
| Lot closes while you are down | Final price never recorded | Usually nobody notices | None: the moment has passed |
The quiet failures cost more than the loud ones. A blocked IP throws errors and wakes someone up. A renamed label that leaves odometer_status empty on a fifth of lots can sit in your database for weeks, and every catalog card and VIN page built on it is wrong in the meantime. Field-level validation (share of lots with a VIN, with photos, with a sale date, with a damage value) is not optional.
The archive you cannot backfill
For most products built on auction data, the valuable part is not the active inventory but the history: what a 2017 Honda Civic with front damage and a salvage title actually sold for, and what this VIN looked like the last two times it went through an auction. That history has one property that changes the whole build-or-buy calculation: it has to be recorded at the time, because final prices and the full lot details are hard to get from the public pages once a sale is over.
So a scraper started today gives you a sold-price archive that starts today. A year of history takes a year. Every outage on a sale day leaves a hole you cannot fill later, and a wrong parse that went unnoticed corrupts history rather than a page you can simply reload. If your product is a price tool, an import calculator or VIN history pages, that is the cost that matters, and no engineering budget shortens it.
What a Copart scraper costs: the structure
We will not publish dollar figures: salaries, proxy prices and server costs vary too much by country and vendor. The structure is stable, though, and most of it is recurring:
| Cost line | Type | What drives it |
|---|---|---|
| Initial build (both auctions) | One-off | Discovery, parsing, normalisation, storage, scheduling |
| Maintenance and fixes | Recurring | How often the sites change; how many sources you cover |
| Proxies / network | Recurring, scales with volume | Requests per hour and how hard the sites push back |
| Servers and browsers | Recurring, scales with volume | Headless browsers are far heavier than plain HTTP |
| Image storage and bandwidth | Recurring, grows forever | Photos per lot times lots per year, if you host them |
| Monitoring and on-call | Recurring | How fast you need to notice a quiet failure |
| Missed history | Permanent | Every day the scraper was down or wrong on a sale day |
A worked example: request volume
Here is a transparent calculation for the network side. Every input below is an assumption for the sake of the example; replace it with your own.
- Assumption: you want 100,000 active lots across Copart and IAAI in your catalog.
- Assumption: you want bids and Buy Now prices no older than one hour, so each active lot is re-read once an hour.
- Assumption: one request per lot (no separate request for photos or a second data call).
That is 100,000 fetches an hour, about 28 a second around the clock, roughly 2.4 million a day, before discovery pages, retries and closing-time checks for final prices. Every one of them goes through the proxy layer and the bot protection described above.
Compare a change feed. Assume (again, an assumption) that a fifth of those lots change in a given hour: 20,000 vehicles. An hourly /cars?minutes=75 request at per_page=1000 returns them in about 20 pages, and /archived-lots?minutes=75 adds a few more for lots that closed. Tens of requests an hour instead of a hundred thousand, because the work of noticing what changed has already been done once, upstream.
A worked example: maintenance hours
Same approach, all inputs are assumptions: two breaking changes a month across both auctions at six hours each (diagnose, fix, re-parse, verify), one hour a week reviewing monitoring and data quality, four hours a month on proxies and infrastructure. That is about 20 engineer hours a month, every month, on top of the initial build. Multiply by your own hourly cost and compare it with a data subscription; then remember that the subscription also comes with history you would otherwise not have.
Copart scraper vs API, side by side
- Days to a prototype, weeks to production
- Sold-price history starts the day you start
- Breaks when page layout or protection changes
- Refresh means re-reading every active lot
- Proxies, browsers and image storage on your bill
- IAAI is a second scraper with its own labels
- Full control over what you collect
- Minutes to the first real record
- Archive with final prices from before you signed up
- The provider fixes breakages
- Change feeds: only what moved since the last run
- One bill, no infrastructure to run
- Copart and IAAI in one schema
- Limited to the fields and sources offered
When building your own scraper is reasonable
We sell data, so weigh this accordingly, but there are cases where we would build it ourselves too:
- A narrow, short-lived need. A one-off research project on a few hundred lots, where history and freshness do not matter.
- A field nobody provides. If your product depends on something no provider extracts, you may have to collect it, and you should ask providers first.
- A site nobody covers. A regional auction or a local marketplace with no data provider. Ask anyway: we add missing sites on request, often within a day.
- You already run collection at scale. If you have the people, the monitoring and the infrastructure for other sources, the marginal cost is lower, though the archive problem remains.
- You are the data provider. Then collection is the product, and you are signing up for all of the above on purpose.
If your situation is a catalog of US salvage auctions, VIN history pages, sold-price tools or an import calculator, the maths usually points the other way: the build is the small part, and the history and maintenance are the large part.
If you build it, monitor these
If you decide to build, budget for monitoring from the first day. These are the checks we would not run a collector without, each one measured per source and per hour:
- Lots discovered per hour against the same hour last week. A sudden drop means discovery broke, not that the auction went quiet.
- Share of lots with each required field: VIN, sale date, damage, odometer, at least one photo. A field that drops from 99% to 80% is a layout change.
- Price sanity: bids and Buy Now within plausible ranges, final bid not below zero, no current bid above Buy Now on a lot still for sale.
- Closed lots without a final price, per sale day. This is the archive leaking.
- Request success rate and latency through the proxy layer, separately for each target site.
- Image 404 rate on a sample of stored URLs, both fresh and a few weeks old.
- Freshness: the age of the newest record per source. An empty error log with stale data is the worst state to be in.
Every check needs an owner who gets the alert and time set aside to act on it. That person's hours are the maintenance line in the cost table above.
What to do next
Before you decide, test the alternative for an afternoon. A free demo key is enough to run the quickstart and see real Copart and IAAI records, including archived lots with final prices. Then read how the hourly sync works in keeping an auction catalog in sync, and if a full site is the goal, how to build a car auction website. The Copart and IAAI API page lists what the feed includes.
Questions people ask
Is it legal to scrape Copart?
It depends on your country, what you collect and how you use it, and Copart's terms of use restrict automated collection from its site. Breaching terms can lead to blocked access or account bans even where scraping itself is not unlawful. This is not legal advice: if your business would depend on scraping, ask a lawyer in your jurisdiction before you build.
How much does it cost to run a Copart scraper?
The initial build is the small part. Recurring costs are maintenance hours when the site changes, proxies and servers that scale with request volume, image storage, and monitoring. As a worked example with stated assumptions, two breakages a month plus routine checks come to about 20 engineer hours a month before infrastructure.
Can I get Copart sold prices by scraping?
Only from the day you start, and only for sales your scraper saw close. Final prices are hard to get from public pages once an auction is over, so history cannot be backfilled later, and every outage on a sale day leaves a permanent gap. A provider that has recorded sales for years gives you that history immediately.
Does Copart block scrapers?
Copart, like other large auction sites, protects its site against automated traffic with rate limits, IP blocking and browser challenges, and changes those measures over time. We do not describe ways around them. For a product, the useful conclusion is that the availability of a scraper is outside your control.
How often should Copart data be refreshed?
For a catalog, hourly is the usual rhythm: bids and Buy Now prices change before the sale, and final prices arrive when it closes. With a change feed you read only what moved since the last run. Polling an API more often than every 10 to 15 minutes rarely brings anything new.


