Fake Addresses

Data sources and licensing

Fake Addresses builds its dataset from public-domain and permissively licensed reference data only, crediting any source whose licence requires attribution exactly as that licence asks. Nothing is scraped, nothing is bought, and no input carries a share-alike obligation. No dataset used here describes a real, identifiable person.

What Kind of Data Fake Addresses Uses

Fake Addresses is built from published reference data: official street and boundary geography, official population statistics, published telephone numbering allocations, and the public postal formatting standards that define how an address in a given country is written. These are the same categories of reference data that mapping and logistics software is built from, and they describe places rather than people.

Every input is either in the public domain or published under a permissive licence. Fake Addresses deliberately avoids share-alike licensed data, because a share-alike obligation would propagate into anything derived from it and restrict what users could do with the output. That constraint rules out some otherwise convenient datasets, and it is applied before a source is adopted rather than after. A dataset that would be convenient but carries an obligation travelling downstream to users is simply not used, and that decision is made once, at the point of adoption, rather than argued about later when something already depends on it.

Nothing is scraped: Fake Addresses does not crawl other address generators, does not lift data from mapping services whose terms forbid it, and does not purchase consumer records. The specific datasets, their versions and their individual licence terms are recorded internally and reviewed whenever a source changes. A source is named on this page only when its own licence requires that credit, and every other input stays unnamed so that publishing a full inventory never becomes a substitute for whether the output is trustworthy. What follows is the part that actually affects a user: what the data is not, what a source's licence requires when one applies, what a user may do with the output, and where the coverage stops.

Sources Credited by Name

Fake Addresses names a source on this page when, and only when, that source's own licence makes attribution a condition of use. Every other input behind the site's published data — everything behind current United States coverage — requires no attribution under its licence and stays unnamed, exactly as before.

Two sources currently meet that bar. The Geocoded National Address File, published by Geoscape Australia via data.gov.au under the Open G-NAF End User Licence Agreement, requires this exact credit wherever the data is shared: "G-NAF © Geoscape Australia licensed by the Commonwealth of Australia under the Open Geo-coded National Address File (G-NAF) End User Licence Agreement."

Code-Point Open and the ONS Postcode Directory, published under the Open Government Licence v3.0, require a three-part credit that must appear together. Its exact required wording is: "Contains OS data © Crown copyright and database right 2026", "Contains Royal Mail data © Royal Mail copyright and database right 2026", and "Source: Office for National Statistics licensed under the Open Government Licence v.3.0."

Neither source has shipped in any address Fake Addresses generates. Both are being ingested for coverage outside the United States, and neither is available on the site today; this page carries their required credit now because attribution is a condition of the licence at the point the data is used or shared, not a courtesy to add once a country page exists. If and when either source ships in generated output, the credit stated above stays exactly as its licence requires — the wording does not change with publication.

No Personal Data Is Involved

None of the reference data behind Fake Addresses describes an identifiable living person. The geography describes streets and administrative boundaries. The population figures are aggregate counts, not records about individuals. The numbering data describes which telephone ranges exist, not who holds a number.

That matters for a specific reason: a generated record cannot accidentally reproduce a real person, because no personal data exists anywhere in the pipeline for it to be drawn from. The names a generated profile carries come from frequency-ranked reference lists, and a given name and a surname are drawn independently, so a produced combination is not a real pairing taken from a source.

Fake Addresses also never emits a value capable of reaching a real person or system. Telephone numbers come from ranges reserved for fiction, email addresses use domains reserved by standard for documentation and testing, and any identifier field uses a pattern that is invalid by construction. Those ranges are chosen so that a test suite which accidentally sends a message, charges a card or submits an identifier cannot reach anybody. A generated record is inert by design, not merely unlikely to belong to someone. That property is asserted on every record the build produces rather than sampled, so a value outside those reserved ranges fails the build instead of reaching a user.

What You Can Do With the Output

Fake Addresses output is intended for testing, development, demonstration and documentation. You can commit it to a repository as test fixtures, ship it inside a demo, paste it into a bug report, or use it to seed a staging database, without tracking down a licence obligation first.

No input licence restricts that use, which is the practical reason the share-alike exclusion above exists. It is also the reason Fake Addresses does not reproduce another generator's output: doing so would inherit whatever terms that generator attaches to its data.

What the output must not be used for is set out in full on the acceptable-use page, and the prohibited-use notice appears on every page that generates an address. The short version is that this data exists to exercise software, not to mislead a person or an institution. Fake Addresses states that on the page itself rather than burying it in linked terms, because a person who needs the warning is exactly the person who will not click through to find it. The distinction that matters is between exercising software, which this data is built for, and presenting a fabricated address to someone entitled to a true one, which it must never be used for.

Limits Worth Knowing

Fake Addresses publishes the limits of its own coverage rather than implying completeness. The dataset is United States only today, and within the United States it covers a defined subset of postal areas and counties rather than the whole country. The exact figures are published by the coverage endpoint of the API and are derived from the dataset itself, never typed by hand.

Telephone area codes are accurate to the state rather than the county, because no freely available source maps numbering ranges to counties with the precision this build would need to claim otherwise. The mailing city attached to a record is currently derived from official place boundaries rather than from the postal service's own preferred city name for that postal code; those two disagree more often than people expect, and Fake Addresses says so rather than implying the stronger guarantee.

Where a limit like this is resolved, the change ships as a new dataset version rather than silently altering the data behind an existing one, so a saved seed keeps returning the records it always did. Publishing a limit costs less than having a user discover it themselves partway through building something on top of it, and every figure quoted anywhere on this site is derived from the dataset rather than written by hand, so a stale claim fails the build instead of reaching a reader.

Frequently asked questions

Which specific datasets does Fake Addresses use?

Fake Addresses does not name most of the datasets behind its coverage. A source that carries no attribution requirement under its own licence — everything behind current United States coverage — stays unnamed, because publishing a full inventory would hand over the method rather than the result. A source whose licence requires attribution is named on this page in that licence's own required wording; two currently meet that bar, both used for coverage outside the United States that has not yet shipped in any published country.

Is Fake Addresses data scraped from other address generators?

No. Nothing in the pipeline is scraped, from a competitor or from anywhere else. Reproducing another generator's output would also inherit whatever licence terms that generator attaches to its data, which is exactly the kind of obligation Fake Addresses avoids by construction. Every input is published reference data that Fake Addresses is entitled to use.

Does any Fake Addresses input licence restrict what I can do with the output?

No Fake Addresses input carries a share-alike obligation, which is the licence type that would otherwise propagate into anything you built from the output. Sources are screened for that before adoption rather than after. If a future source ever did impose an obligation on downstream use, Fake Addresses would state it plainly on this page rather than leave a user to discover it.

Are the credited Australian and UK sources live on the site yet?

Fake Addresses is United States only today, and neither source appears in any address the site generates. The Geocoded National Address File and Code-Point Open, together with the ONS Postcode Directory, are credited on this page because their own licences require it, not because their data is live. Neither Australia nor the United Kingdom is a published country yet.