Skip to content

Copart scraper or API: what scraping Copart really costs

What a Copart scraper really needs: proxies, layout fixes, images, monitoring and a sold archive you cannot backfill. The cost structure, and when to build.

9 min read
Illustration: a fragile scraper pipeline with proxy, parser and image boxes, some marked broken, next to a single API request returning clean JSON

A Copart scraper is cheap to write and expensive to keep. A developer with a headless browser can pull lot pages into JSON in a day or two; the cost arrives afterwards, as proxies, bot protection, layout changes, image handling, monitoring and the hours spent fixing it, every month, for as long as your product depends on it. And the one thing money cannot buy later, an archive of sold lots with final prices, only exists from the day your scraper started.

We have collected Copart and IAAI data every day for years, so this is written from the maintenance side. It is not a how-to for getting around anyone's defences: we do not describe that here. It is a cost breakdown, with the assumptions written out so you can plug in your own numbers, and an honest list of cases where building your own collector is the right decision.

Key takeaways

  • The first version of a Copart scraper is the cheap part; maintenance is a recurring cost with no end date.
  • Bot protection, layout changes and image URLs are the usual breaking points, and you notice them only if you monitor for them.
  • Sold prices have to be captured when each auction closes; a scraper that starts today has no history before today, and every outage leaves a permanent gap.
  • Refreshing a large catalog by re-reading pages costs orders of magnitude more requests than reading a change feed.
  • Building your own is reasonable for narrow, short-lived or unusual needs, and for sites no provider covers.

What a Copart scraper has to do

"Scrape Copart" sounds like one job. In production it is six, and each one fails in its own way.

  1. 1DiscoverFind every lot: search pages, sale lists, new arrivals. Miss a page and those lots never exist for you.
  2. 2FetchLoad each lot page or the data behind it, within rate limits, through proxies, with retries.
  3. 3ParseTurn labels into fields: VIN, odometer and its status, damage, title, condition, bids, dates, location.
  4. 4NormaliseMap wording to stable values, convert times to UTC, split vehicle from lot, dedupe relists.
  5. 5Track changesRe-read active lots for bids and Buy Now, notice when a lot closes, capture the final price.
  6. 6OperateStore photos or links, monitor every step, alert, fix, redeploy. Forever.
The pipeline behind a working scraper. The first three steps are the demo; the last three are the product.

Most teams estimate the first three steps and discover the other three in production. Normalisation alone is a project if you want IAAI next to Copart: different labels, different identifiers, different page structure (the Copart vs IAAI comparison lists them).

Bot protection, proxies and terms of use

Large auction sites protect themselves against automated traffic, and Copart and IAAI are no exception. In practice a scraper at catalog scale meets rate limits, blocked IP ranges and browser challenges, and the rules change without notice. Teams respond with proxy pools and browser automation, which turns into a running cost and an arms race whose pace someone else sets. We are not going to describe how to defeat any of those measures, and we would not build a business on doing so.

There is also the contract question. Both auctions' terms of use restrict automated collection from their sites, and the legal position of scraping differs by country and depends on what you collect and how you use it. That is a conversation for your lawyer, not for an engineering blog, but it belongs in the cost column: a plan whose central dependency can be withdrawn at any time is a risk you are pricing in, whether you write it down or not.

Where Copart scrapers break

Here are the incidents we plan for, roughly in the order a new scraper meets them. None of them is exotic; together they are why "it worked last week" is the most common sentence in scraper maintenance.

IncidentWhat you seeHow you noticeTypical fix
Page layout or data structure changesEmpty fields, wrong values in the right fieldsOnly with field-level validationRewrite selectors, re-parse affected lots
IP range blocked or challengedError pages, timeouts, sudden zero resultsSuccess-rate alertsInfrastructure work, more proxy spend
Site maintenance at nightGaps in a runGap detection per hourRe-run the window, if the data is still there
New price format or currency displayParsing errors or silently wrong numbersRange checks on pricesParser fix plus data repair
New field appears (e.g. a new badge)Nothing: you just do not have itSomeone compares with the site by handSchema change, backfill impossible
Image URLs change or expireBroken photos on your siteImage 404 rateRe-fetch or re-host media
Lot closes while you are downFinal price never recordedUsually nobody noticesNone: the moment has passed
Common scraper failures. The last row is the expensive one.

The quiet failures cost more than the loud ones. A blocked IP throws errors and wakes someone up. A renamed label that leaves odometer_status empty on a fifth of lots can sit in your database for weeks, and every catalog card and VIN page built on it is wrong in the meantime. Field-level validation (share of lots with a VIN, with photos, with a sale date, with a damage value) is not optional.

The archive you cannot backfill

For most products built on auction data, the valuable part is not the active inventory but the history: what a 2017 Honda Civic with front damage and a salvage title actually sold for, and what this VIN looked like the last two times it went through an auction. That history has one property that changes the whole build-or-buy calculation: it has to be recorded at the time, because final prices and the full lot details are hard to get from the public pages once a sale is over.

So a scraper started today gives you a sold-price archive that starts today. A year of history takes a year. Every outage on a sale day leaves a hole you cannot fill later, and a wrong parse that went unnoticed corrupts history rather than a page you can simply reload. If your product is a price tool, an import calculator or VIN history pages, that is the cost that matters, and no engineering budget shortens it.

What a Copart scraper costs: the structure

We will not publish dollar figures: salaries, proxy prices and server costs vary too much by country and vendor. The structure is stable, though, and most of it is recurring:

Cost lineTypeWhat drives it
Initial build (both auctions)One-offDiscovery, parsing, normalisation, storage, scheduling
Maintenance and fixesRecurringHow often the sites change; how many sources you cover
Proxies / networkRecurring, scales with volumeRequests per hour and how hard the sites push back
Servers and browsersRecurring, scales with volumeHeadless browsers are far heavier than plain HTTP
Image storage and bandwidthRecurring, grows foreverPhotos per lot times lots per year, if you host them
Monitoring and on-callRecurringHow fast you need to notice a quiet failure
Missed historyPermanentEvery day the scraper was down or wrong on a sale day
Cost lines of an in-house scraper. Only the first one ends.

A worked example: request volume

Here is a transparent calculation for the network side. Every input below is an assumption for the sake of the example; replace it with your own.

  • Assumption: you want 100,000 active lots across Copart and IAAI in your catalog.
  • Assumption: you want bids and Buy Now prices no older than one hour, so each active lot is re-read once an hour.
  • Assumption: one request per lot (no separate request for photos or a second data call).

That is 100,000 fetches an hour, about 28 a second around the clock, roughly 2.4 million a day, before discovery pages, retries and closing-time checks for final prices. Every one of them goes through the proxy layer and the bot protection described above.

Compare a change feed. Assume (again, an assumption) that a fifth of those lots change in a given hour: 20,000 vehicles. An hourly /cars?minutes=75 request at per_page=1000 returns them in about 20 pages, and /archived-lots?minutes=75 adds a few more for lots that closed. Tens of requests an hour instead of a hundred thousand, because the work of noticing what changed has already been done once, upstream.

A worked example: maintenance hours

Same approach, all inputs are assumptions: two breaking changes a month across both auctions at six hours each (diagnose, fix, re-parse, verify), one hour a week reviewing monitoring and data quality, four hours a month on proxies and infrastructure. That is about 20 engineer hours a month, every month, on top of the initial build. Multiply by your own hourly cost and compare it with a data subscription; then remember that the subscription also comes with history you would otherwise not have.

Copart scraper vs API, side by side

Your own scraper
  • Days to a prototype, weeks to production
  • Sold-price history starts the day you start
  • Breaks when page layout or protection changes
  • Refresh means re-reading every active lot
  • Proxies, browsers and image storage on your bill
  • IAAI is a second scraper with its own labels
  • Full control over what you collect
Auction data API
  • Minutes to the first real record
  • Archive with final prices from before you signed up
  • The provider fixes breakages
  • Change feeds: only what moved since the last run
  • One bill, no infrastructure to run
  • Copart and IAAI in one schema
  • Limited to the fields and sources offered
The trade-off in one picture. The last line on each side is the real argument.

When building your own scraper is reasonable

We sell data, so weigh this accordingly, but there are cases where we would build it ourselves too:

  • A narrow, short-lived need. A one-off research project on a few hundred lots, where history and freshness do not matter.
  • A field nobody provides. If your product depends on something no provider extracts, you may have to collect it, and you should ask providers first.
  • A site nobody covers. A regional auction or a local marketplace with no data provider. Ask anyway: we add missing sites on request, often within a day.
  • You already run collection at scale. If you have the people, the monitoring and the infrastructure for other sources, the marginal cost is lower, though the archive problem remains.
  • You are the data provider. Then collection is the product, and you are signing up for all of the above on purpose.

If your situation is a catalog of US salvage auctions, VIN history pages, sold-price tools or an import calculator, the maths usually points the other way: the build is the small part, and the history and maintenance are the large part.

If you build it, monitor these

If you decide to build, budget for monitoring from the first day. These are the checks we would not run a collector without, each one measured per source and per hour:

  1. Lots discovered per hour against the same hour last week. A sudden drop means discovery broke, not that the auction went quiet.
  2. Share of lots with each required field: VIN, sale date, damage, odometer, at least one photo. A field that drops from 99% to 80% is a layout change.
  3. Price sanity: bids and Buy Now within plausible ranges, final bid not below zero, no current bid above Buy Now on a lot still for sale.
  4. Closed lots without a final price, per sale day. This is the archive leaking.
  5. Request success rate and latency through the proxy layer, separately for each target site.
  6. Image 404 rate on a sample of stored URLs, both fresh and a few weeks old.
  7. Freshness: the age of the newest record per source. An empty error log with stale data is the worst state to be in.

Every check needs an owner who gets the alert and time set aside to act on it. That person's hours are the maintenance line in the cost table above.

What to do next

Before you decide, test the alternative for an afternoon. A free demo key is enough to run the quickstart and see real Copart and IAAI records, including archived lots with final prices. Then read how the hourly sync works in keeping an auction catalog in sync, and if a full site is the goal, how to build a car auction website. The Copart and IAAI API page lists what the feed includes.

Questions people ask

Is it legal to scrape Copart?

It depends on your country, what you collect and how you use it, and Copart's terms of use restrict automated collection from its site. Breaching terms can lead to blocked access or account bans even where scraping itself is not unlawful. This is not legal advice: if your business would depend on scraping, ask a lawyer in your jurisdiction before you build.

How much does it cost to run a Copart scraper?

The initial build is the small part. Recurring costs are maintenance hours when the site changes, proxies and servers that scale with request volume, image storage, and monitoring. As a worked example with stated assumptions, two breakages a month plus routine checks come to about 20 engineer hours a month before infrastructure.

Can I get Copart sold prices by scraping?

Only from the day you start, and only for sales your scraper saw close. Final prices are hard to get from public pages once an auction is over, so history cannot be backfilled later, and every outage on a sale day leaves a permanent gap. A provider that has recorded sales for years gives you that history immediately.

Does Copart block scrapers?

Copart, like other large auction sites, protects its site against automated traffic with rate limits, IP blocking and browser challenges, and changes those measures over time. We do not describe ways around them. For a product, the useful conclusion is that the availability of a scraper is outside your control.

How often should Copart data be refreshed?

For a catalog, hourly is the usual rhythm: bids and Buy Now prices change before the sale, and final prices arrive when it closes. With a change feed you read only what moved since the last run. Polling an API more often than every 10 to 15 minutes rarely brings anything new.

© 2025. AuctionsAPI operates independently and is not affiliated with Copart, IAAI or Encar.