~ / guides / Where Does Zillow Get Its Data? (ATTOM, MLS & GreatSchools)

Where Does Zillow Get Its Data? (ATTOM, MLS & GreatSchools)

DC
Dana Cole
Zillow data engineer · about the author
the short version
  • Zillow does not own most of its data. A single listing page stitches together MLS IDX feeds (active listings), county records (tax, deeds, assessments), and several third-party licensors for schools, flood risk, and parcel maps.
  • Listings come from MLS Internet Data Exchange (IDX) feeds. Zillow switched to direct IDX in 2021 after joining MLSs as a dues-paying broker member.
  • School ratings come from GreatSchools, which licenses the same ratings to Zillow, Redfin, and Realtor.com. Tax and price history come from public county records aggregated by a third-party data provider.
  • On the ATTOM question: there is no public ATTOM-Zillow partnership. ATTOM is a competing aggregator selling the same county-record data Zillow ingests separately.

I spent an afternoon trying to answer one question for myself: where does Zillow get its data? I had been parsing Zillow property pages for years, and I realized I could name the HTML fields better than I could name their origins. So I took a single listing and traced every block on it back to a source: the price history, the tax record, the school rating, the flood score, the “Zestimate,” the comparable sales. None of it originates at Zillow. The company is a real estate data aggregator that licenses, ingests, and stitches together feeds from MLSs, county governments, and a short list of third-party data vendors.

This article maps each Zillow data source to the part of the listing it powers. I will answer the ATTOM question directly, explain how Zillow gets MLS data, and cover the school, flood, and crime data sources that people ask about most.

Quick answer: what is the source data for Zillow?

The source data for Zillow comes from four layers, and a single listing page combines all of them. There is no one Zillow database. The company assembles each property page in real time from feeds it licenses or pulls from public records.

Data on the listingPrimary sourceType of source
Active for-sale / for-rent listingsMLS IDX feedsLicensed feed (real estate)
Price & tax historyCounty records, via a third-party aggregatorPublic records
Property facts (beds, baths, sq ft, parcel)County assessor + MLS + owner-submittedPublic records + user data
Assessment & deed dataCounty recorder / assessor officesPublic records
School ratingsGreatSchools.orgLicensed third-party data
School attendance zonesThird-party location-data providerLicensed third-party data
Flood & climate riskFirst StreetLicensed third-party data
Comparable sales (comps)Recent area sales in public records + MLSDerived
ZestimateZillow algorithm over all of the aboveComputed

Two patterns explain almost everything on the page. Real estate listing fields come from the MLS through IDX feeds. Everything historical or factual about the building (taxes, deeds, square footage, ownership) traces to county government records, which Zillow obtains through a third-party data provider that collects and standardizes them. The remaining specialty layers (schools, flood) are licensed from named vendors. The sections below take each layer in turn, starting with listings, which is the data most people mean when they ask the question.

How does Zillow get its listing data?

Zillow gets its active listing data from MLS Internet Data Exchange (IDX) feeds, the standardized feeds that Multiple Listing Services offer directly to their members. When an agent enters a property into the local MLS, the IDX feed pushes that listing, with its price, photos, and status, to consumer search portals automatically. The same mechanism populates Realtor.com, Redfin, Trulia, and hundreds of individual brokerage sites.

Zillow did not always collect listings this way. In 2021 it overhauled its listing pipeline, moving from thousands of disparate syndication feeds to direct MLS IDX feeds. The CRMLS blog documented the transition, explaining that Zillow joined MLSs across the country to switch to these feeds. The switch was possible because of a structural change: Zillow transitioned from being a listing-distribution partner to a dues-paying member of the National Association of Realtors and local Associations, which gave it the same right to access IDX feeds that any brokerage holds. METRO MLS noted that as a broker member, Zillow began receiving the IDX feed on the same terms as other brokerages.

IDX has limits, and Zillow supplements it. IDX rules govern what can be displayed and for how long, so for sold and off-market history Zillow leans on public records instead of the feed. The company also rolled out a Virtual Office Website (VOW) feed in some markets to supplement IDX, since VOW feeds permit display of data that IDX restricts. For-sale-by-owner listings sit outside the MLS entirely: when a home is sold without an agent, the owner posts the listing manually through a Zillow account, so that data is user-generated and never touches an MLS feed.

That covers what is currently on the market. The deeper history attached to every address, the part that persists after a listing closes, comes from a different source entirely.

Where does Zillow get price, tax, and assessment data?

Zillow gets price history, tax history, and assessment data from public county records, collected and aggregated by a third-party data provider before it reaches Zillow. Every county in the US maintains records of real estate transactions and tax assessments through the office responsible for recording deeds and assessing property. That information is public. Zillow does not visit each county itself: it buys a standardized national feed from a data vendor that collects, cleans, and aggregates those records, then delivers them to Zillow.

This is the layer where the ATTOM question actually lives, so it is worth being precise. The “third-party data provider” Zillow references in its help documentation is the same role ATTOM plays for its own customers. The data category is identical: tax, deed, assessment, and ownership records sourced from county governments. The sourcing path explains why a Zillow tax history and an ATTOM tax history for the same parcel usually agree. Both trace to the same county recorder.

Zillow also built its own historical version of this data. For years it published the Zillow Transaction and Assessment Database (ZTRAX), a research dataset of more than 100 million parcels covering decades of deeds and assessments. Zillow stopped distributing ZTRAX in 2023. The University of Michigan’s ICPSR archive now houses ZTRAX and distributes it to member institutions, with Zillow generating a refreshed dataset twice a year. ZTRAX matters here because it confirms the underlying source in Zillow’s own words: it described ZTRAX as built from county assessor and recorder data, the same public records that feed the live tax and price history you see on a listing today.

Property facts (bedrooms, bathrooms, square footage, lot size, year built) come from this same county-assessor layer, supplemented by the MLS listing and by corrections that owners submit through their Zillow accounts. When those three disagree, which they often do, the listing can show stale or conflicting numbers. The parcel geometry and lot boundaries draw on geographic information system (GIS) data, the spatial parcel records that county and municipal GIS departments maintain.

With the building’s facts and history established, the next question is the headline number Zillow computes on top of them.

Does Zillow provide comparable sales data, and how is the Zestimate built?

Zillow does provide comparable sales data, and comps are one of the inputs to the Zestimate. Comparable sales are recent transactions of similar nearby homes, matched on location, beds, baths, condition, and square footage. Per Zillow’s documentation on real estate comps, it usually surfaces up to five comps for a property, each carrying its own sale price, sale date, and square footage drawn from recent area sales in public records and the MLS.

The Zestimate is a computed figure. It is the output of a Zillow algorithm that runs over the source layers already described, so it inherits whatever its inputs say and originates no data of its own. According to Zillow’s Zestimate documentation, the model ingests public property records, tax assessments, prior sales, MLS on-market data such as list price and days on market, comparable homes, and user-submitted home facts, then applies machine learning to weight those inputs by local market. Zillow reports a nationwide median error rate of roughly 1.9% for on-market homes and about 6.9% for off-market homes, which reflects how much the estimate depends on having a live MLS listing in the mix.

The accuracy gap is the practical takeaway for anyone collecting this data. An on-market Zestimate is anchored by a current MLS listing. An off-market one leans almost entirely on public records and comps, so its error band is roughly three to four times wider. The estimate is only as current as the feeds underneath it. The same dependence on outside licensors shows up in the specialty data layers, schools and flood risk, which Zillow does not compute at all.

Where does Zillow get school data?

Zillow gets its school ratings from GreatSchools.org, a nonprofit that scores K-12 schools and licenses those ratings to real estate companies. The GreatSchools rating you see on a Zillow listing is the same rating that appears on Redfin and Realtor.com, because GreatSchools licenses its K-12 data to real estate, media, and technology companies across the board, with no single portal holding it exclusively. GreatSchools refreshes the school directory and performance data it shares with real estate partners on a regular cycle.

Zillow’s relationship with GreatSchools has been formalized. Zillow Group announced GreatSchools as its school-data partner, with GreatSchools supplying school ratings and profile information across Zillow’s listings. The rating itself blends measures such as test scores, academic progress, and equity indicators, which GreatSchools publishes its methodology for.

School attendance zones are a separate data source from the ratings, and this trips people up. The GreatSchools rating tells you how a school scores. The attendance boundary tells you which school an address is assigned to, and that boundary data has historically come from a third-party location-data provider outside the GreatSchools relationship. The GreatSchools help desk confirms that boundary errors on Zillow originate outside GreatSchools, which is why a listing can show a correct rating attached to the wrong assigned school. Anyone scraping school fields should treat rating and assignment as two distinct, separately sourced values.

Schools are the most stable specialty layer. The flood and climate data layer has been the most volatile, and it changed materially in the last year.

Where does Zillow get its flood and crime data?

Zillow gets its flood and climate risk data from First Street, the climate-analytics company behind the Flood Factor score. First Street’s Flood Factor assigns every US property a 1-to-10 score modeling cumulative flood risk over a 30-year horizon, drawing on rainfall, river and stream overflow, tides, and storm surge. First Street’s flood-risk dataset has been archived publicly for research on Zenodo, which documents its property-level methodology.

The flood data story has a recent twist that directly answers what is on Zillow today. Zillow added First Street climate risk scores to for-sale listings in September 2024, covering flood, wildfire, wind, heat, and air-quality risk, citing that more than 80% of buyers consider climate risk. Then it pulled the on-listing scores back. After the California Regional MLS questioned the accuracy of the scores and agents reported that the data was hurting sales, Zillow removed the climate risk display in late 2025 and replaced the prominent scores with a quieter link out to First Street’s own records. So a property page may now show a First Street link instead of an embedded flood score, depending on when you look.

Crime data is the one major category Zillow does not show. In 2021, Zillow-owned Trulia, along with Redfin and Realtor.com, announced they would stop displaying crime data over concerns about how the maps were sourced and how they could reinforce bias. If you are looking for crime data on a Zillow listing, you will not find a native field for it. That gap is a deliberate sourcing decision the portals made together.

How does this compare to Redfin and Realtor.com?

Zillow, Redfin, and Realtor.com all sit on the same underlying data sources, which is why their listings look so similar. All three pull active listings from MLS IDX feeds, all three license school ratings from GreatSchools, and all three obtain tax and assessment history from public county records. The differences are in the proprietary layers each company builds on top: Zillow’s Zestimate, Redfin’s Redfin Estimate, and each company’s own technology stack for ranking, search, and lead routing.

LayerZillowRedfinRealtor.com
Active listingsMLS IDX feedsMLS IDX + direct (brokerage)MLS feeds (operated under News Corp / Move)
School ratingsGreatSchoolsGreatSchoolsGreatSchools
Tax / assessmentCounty records via aggregatorCounty recordsCounty records
Home value estimateZestimateRedfin EstimateRealEstimate (multiple providers)
Flood / climate riskFirst Street (link as of late 2025)First StreetFirst Street

The shared foundation is the reason a product built to read one portal can usually be adapted to another. The real estate data itself is largely common. What differs is each company’s presentation and its computed estimates. For development teams, that means the hard part is rarely the schema. It is getting reliable access to the listing pages at all, which is where I spend most of my own time.

How do you get structured data out of Zillow?

The realistic way to get structured data out of Zillow is to read the rendered listing pages, because Zillow does not offer a public, general-purpose data API for the property and listing data described above. The fields are all on the page (price, beds, baths, tax history, the Zestimate, school ratings, the comps list), but Zillow defends those pages with bot detection, so a naive request is the part that breaks. I cover the blocking specifics in my guide on scraping Zillow without getting blocked, and the full extraction walkthrough in how to scrape Zillow data.

In my own runs I route Zillow property pages through ChocoData’s API, which takes a Zillow URL and returns the parsed fields as JSON. Here is the request shape I use:

curl "https://chocodata.com/api/v1/zillow/property?url=https://www.zillow.com/homedetails/2092-zpid/&api_key=$CHOCO_API_KEY"

The response gives back the structured property record (address, price, beds and baths, square footage, tax and price history, and the Zestimate) without my having to manage proxies or parse the page myself. For the listing and property fields specifically, the Zillow Listings & Property Data Scraper is the endpoint I reach for. When I need pricing and sales figures in bulk, I use the Zillow Home Price & Sales Data API instead, and for rentals the Zillow Rental Data API. You can sign up for a ChocoData key and run the call above against a live property to see the parsed shape.

Knowing where each field originates is what makes the parsed output useful. A tax-history field traces to county records and updates slowly. A list price comes from the MLS IDX feed and changes the moment an agent edits it. A school rating is a licensed GreatSchools value that refreshes on GreatSchools’ schedule. The data is only ever as fresh and as accurate as the upstream source it came from, and now you know what those sources are.

Before you build a pipeline on top of any of it, it is worth understanding the rules around collecting it, which I cover in is scraping Zillow legal.

FAQ

Does Zillow use ATTOM data?

There is no public record of a Zillow-ATTOM data partnership or customer relationship. ATTOM is a property-data aggregator that licenses tax, deed, and assessment records, the same category of public county data Zillow ingests. Both companies collect from the same underlying government sources, which is why their property facts often match, but Zillow has not named ATTOM as a supplier in its help documentation.

Does Zillow use MLS data?

Yes. Active for-sale and for-rent listings on Zillow come from MLS Internet Data Exchange (IDX) feeds. Zillow moved to direct IDX feeds in 2021 after becoming a dues-paying member of the National Association of Realtors and local MLSs, which gave it the same feed access as any brokerage.

Does Zillow provide comparable sales data?

Yes. Zillow shows recent comparable sales (comps) on listing pages and feeds them into the Zestimate. According to Zillow's own documentation, it usually includes up to five comps drawn from recent area sales in public records and the MLS, each with sale price, sale date, and square footage.

Where does Zillow get its flood data?

Zillow sourced flood and other climate risk scores from First Street starting in September 2024. After the California Regional MLS questioned the scores' accuracy, Zillow removed the on-listing climate risk display in late 2025 and replaced it with a link out to First Street's own records.

How is school data determined on Zillow?

School ratings on Zillow are licensed from GreatSchools.org, which scores schools on test scores, equity, and other measures. School attendance-zone boundaries are a separate dataset. Zillow has historically sourced those boundaries from a third-party location-data provider outside the GreatSchools relationship.

DC
Dana Cole
I've built Zillow data pipelines for years. On zillowscraperapi.com I run Zillow scraping methods against live pages and publish what actually holds up.