snowfakery
Snowfakery is a tool for generating fake data that has relations between tables. Every row is faked data, but also unique and random, like a snowflake.
Decision gist · record as of 2026-08-14
Yes. Snowfakery is actively maintained, has no known vulnerabilities, installs with low friction, and solves a real problem—generating realistic relational test data from declarative recipes. It is particularly valuable if you work with Salesforce via CumulusCI or need repeatable synthetic datasets for testing. The BSD license imposes no practical restrictions.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.11 or later.
- Output to Salesforce requires CumulusCI integration.
- Low friction install with a pure-Python wheel and 13 runtime dependencies.
License · maintenance · safety
BSD 3-Clause License (permissive) — BSD 3-Clause License is permissive; you can use, modify, and distribute this package with minimal restrictions, provided you include the license notice.
last release 2026-01-09 (217 days) · last repo commit 2026-04-27 · 157 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 237,104 downloads/mo, #8,966 on PyPI
Alternatives
Verify before relying
pip install snowfakery
snowfakery generate recipe.yaml
Or in Python:
from snowfakery import generate_data
generate_data(recipe_file='recipe.yaml')- Whether the package can generate data to custom output formats beyond stdout and SQLAlchemy databases without code modification
- Performance characteristics when generating large datasets or deeply nested relational structures
What it is and what it does
Snowfakery is a data generation tool designed to create fake but realistic relational datasets from declarative YAML recipes. It generates unique, random rows across related tables while maintaining referential integrity—useful for testing, development, and populating Salesforce orgs. The tool depends on faker for base data generation, SQLAlchemy for database connectivity, and Jinja2 for template rendering within recipes.
You define what data to generate in a YAML recipe file, specifying tables, relationships, and field patterns. Snowfakery then produces output to stdout, any SQLAlchemy-compatible database, or directly to Salesforce when run through CumulusCI. The package is actively maintained, supports Python 3.11–3.13, and carries no known security vulnerabilities.
Use it for
- Populate test databases with realistic relational data for integration tests without hardcoding fixtures
- Generate sample Salesforce orgs with linked records for UAT or demo environments
- Create anonymized or synthetic datasets for development when production data cannot be used
- Seed databases with consistent, reproducible fake data across multiple test runs
- Rapidly prototype data models by generating sample data matching a schema
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Snowfakery is actively maintained, has no known vulnerabilities, installs with low friction, and solves a real problem—generating realistic relational test data from declarative recipes. It is particularly valuable if you work with Salesforce via CumulusCI or need repeatable synthetic datasets for testing. The BSD license imposes no practical restrictions.
Install
snowfakery on PyPI
Before you install
Low friction install with a pure-Python wheel and 13 runtime dependencies. Actively maintained with a recent commit on 2026-04-27 and no known vulnerabilities.
Requires Python 3.11 or later. Output to Salesforce requires CumulusCI integration.
License in practice
BSD 3-Clause License is permissive; you can use, modify, and distribute this package with minimal restrictions, provided you include the license notice.
Quickstart
pip install snowfakery
snowfakery generate recipe.yaml
Or in Python:
from snowfakery import generate_data
generate_data(recipe_file='recipe.yaml')
Verify before relying
- Whether the package can generate data to custom output formats beyond stdout and SQLAlchemy databases without code modification
- Performance characteristics when generating large datasets or deeply nested relational structures
Package facts
| License | BSD 3-Clause License permissive |
| Python support | Supports the current Python release >=3.11 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 13 packagesclickfakerfaker-edufaker-nonprofitgvgenjinja2pydanticpython-baseconvpython-dateutilpyyamlrequestssetuptoolssqlalchemy |
| Maintenance | Actively maintained 217 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 237,104 / month, #8,966 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: BSD LicenseNatural Language :: EnglishProgramming Language :: Python :: 3Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13 |
Evidence: snowfakery-4.2.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “fake data generation with relations”
- snowfakerySnowfakery generates fake relational data from YAML recipes, writing…
- FakerFaker generates realistic fake data—names, addresses, emails, phone…
- jsfGenerates realistic fake JSON data from JSON Schema definitions,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Testing packages
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.
Install it if you write tests for any asyncio-based code.
A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.
Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.
See also faker-nonprofit · cumulusci · faker-edu · fakesnow · streamlit-faker · Faker · robotframework-faker · mimesis · fake-factory · jsf