names-dataset
The python library to handle names
Decision gist · record as of 2026-08-14
Yes, if you need name-to-demographics lookup and can accept the 3.2GB memory cost and aging codebase. The package is stable (no known vulnerabilities, permissive license), but expect no active maintenance—use it for read-only demographic queries, not as a foundation for a critical service. Low install friction and reasonable download volume suggest it works in practice for its intended use.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires 3.2GB of RAM to load the full dataset in memory; initialization is slow and should be done once at application startup.
- Low install friction with just two runtime dependencies (pycountry, numpy).
- Package is aging—last release was 493 days ago—so expect no active maintenance or bug fixes.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; the underlying data derives from a public Facebook leak, and names themselves are generally not copyrightable, though legal review is recommended for sensitive use cases.
last release 2025-04-08 (493 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 112,862 downloads/mo, #12,356 on PyPI
Alternatives
Verify before relying
pip install names-dataset
from names_dataset import NameDataset, NameWrapper
nd = NameDataset()
result = nd.search('Philippe')
print(NameWrapper(result).describe) # Male, France- Whether the 105 countries claim in the summary matches the 106 countries mentioned in the full dataset section.
- Current state of the GitHub repository (archived status, commit history) since maintenance.repo fields are null.
- Whether Python version support is truly unspecified or if there are undocumented constraints.
What it is and what it does
names-dataset is a lookup library that maps first and last names to demographic attributes: the countries where they are most common (with probability distributions), gender likelihood, and popularity rank within each country. It wraps a dataset of 730K first names and 983K last names extracted from a public Facebook data leak, covering 105 countries. The library loads this data into memory at startup and provides search, ranking, and autocomplete/fuzzy-match APIs.
You use it to answer questions like "Is Philippe more likely male or female, and from which country?" or to retrieve the top 10 most popular male names in the United States. It supports both exact search and fuzzy matching (to handle misspellings) and real-time autocomplete. The main constraint is its large memory footprint—3.2GB—which means it's best suited for applications where you initialize it once and reuse it, rather than spinning up fresh instances frequently.
Use it for
- Predict gender and likely country of origin from a person's name in user registration or data-cleaning workflows.
- Populate autocomplete dropdowns in forms that ask for first or last names, with real-time suggestions as the user types.
- Rank names by popularity within a specific country to identify common vs. rare names for analysis or validation.
- Correct misspelled names via fuzzy search, e.g., matching 'Isablel' to 'Isabel' in data import pipelines.
- Generate synthetic or representative name lists for testing, localization, or demographic analysis across multiple countries.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need name-to-demographics lookup and can accept the 3.2GB memory cost and aging codebase.
The package is stable (no known vulnerabilities, permissive license), but expect no active maintenance—use it for read-only demographic queries, not as a foundation for a critical service. Low install friction and reasonable download volume suggest it works in practice for its intended use.
Install
names-dataset on PyPI
Before you install
Low install friction with just two runtime dependencies (pycountry, numpy). Package is aging—last release was 493 days ago—so expect no active maintenance or bug fixes.
Requires 3.2GB of RAM to load the full dataset in memory; initialization is slow and should be done once at application startup.
License in practice
MIT license permits commercial and private use with minimal restrictions; the underlying data derives from a public Facebook leak, and names themselves are generally not copyrightable, though legal review is recommended for sensitive use cases.
Quickstart
pip install names-dataset
from names_dataset import NameDataset, NameWrapper
nd = NameDataset()
result = nd.search('Philippe')
print(NameWrapper(result).describe) # Male, France
Verify before relying
- Whether the 105 countries claim in the summary matches the 106 countries mentioned in the full dataset section.
- Current state of the GitHub repository (archived status, commit history) since maintenance.repo fields are null.
- Whether Python version support is truly unspecified or if there are undocumented constraints.
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagespycountrynumpy |
| Maintenance | Aging 493 days since the last release |
| First released | |
| Downloads | 112,862 / month, #12,356 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: names_dataset-3.3.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “name demographics lookup”
- names-datasetLooks up demographic information about first and last names across…
- gender-guesserGuesses the gender of a person from their first name using a lookup…
- GenderizeGenderize is a client library that queries the Genderize.io web…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also gender-guesser · nicknames · Genderize · country_list · names · country-converter · geonamescache · countryinfo · getname · python-jobspy