dbl-tempo
Tempo is timeseries manipulation for Spark. This project builds upon the capabilities of PySpark to provide a suite of abstractions and functions that make operations on timeseries data easier and highly scalable.
Decision gist · record as of 2026-08-14
Yes. Tempo is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and zero runtime dependencies. It is well-suited for teams already using Databricks who need to simplify time series operations. Install it if you work with time-indexed data; skip it if you do not need time series transformations.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Apache Spark and PySpark; designed for use in Databricks notebooks or local Spark environments.
- Low install friction with a pure-Python wheel.
- Actively maintained as of 2026-07-10 with 344 repository stars and regular releases since 2021.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, making it suitable for most production environments.
last release 2025-09-15 (333 days) · last repo commit 2026-07-10 · 344 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 4,221,234 downloads/mo, #2,360 on PyPI
Alternatives
Verify before relying
pip install dbl-tempo
from tempo import TSDF
tsdf = TSDF(spark_df, ts_col="event_ts", partition_cols=["user_id"])
resampled = tsdf.resample(freq='min', func='mean')- Whether Tempo's operations scale efficiently to very large time series datasets or specific volume thresholds.
- Performance characteristics of AS OF joins compared to native Spark windowing on typical workloads.
- Support for custom aggregation functions beyond the documented 'floor', 'ceil', 'min', 'max', 'mean' options.
What it is and what it does
Tempo is a library that wraps DataFrames to simplify time series analysis. It provides a TSDF (Time Series DataFrame) abstraction that treats a DataFrame as a collection of time series, one per partition key, and exposes methods for common time series operations: resampling to different frequencies, AS OF joins to merge the latest records from a source table, rolling statistics over time windows, exponential and simple moving averages, Fourier transforms, and interpolation of missing values. The library also supports Delta Lake optimization on time and partition fields.
Tempo is built for data teams using Databricks who need to perform time series transformations at scale. It handles the complexity of windowing and sorting across partitions, allowing you to work with time-indexed data using a higher-level API. The entry point is always a TSDF object, which requires a distinguished timestamp column and optionally partition and sequence columns.
Use it for
- Resample high-frequency sensor or event data into regular time buckets (e.g., 1-minute aggregates) for analysis or visualization.
- Perform AS OF joins to enrich a fact table with the most recent values from a slowly-changing dimension or reference table.
- Generate lagged features from time series for machine learning (e.g., prior 50 values for EMA-based prediction).
- Compute rolling statistics (mean, min, max) over a time window to detect anomalies or trends in streaming data.
- Interpolate missing time periods in sparse time series using forward-fill, backward-fill, or linear interpolation methods.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Tempo is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and zero runtime dependencies. It is well-suited for teams already using Databricks who need to simplify time series operations. Install it if you work with time-indexed data; skip it if you do not need time series transformations.
Install
dbl-tempo on PyPI
Before you install
Low install friction with a pure-Python wheel. Actively maintained as of 2026-07-10 with 344 repository stars and regular releases since 2021.
Requires Apache Spark and PySpark; designed for use in Databricks notebooks or local Spark environments.
License in practice
MIT license permits commercial and private use with minimal restrictions, making it suitable for most production environments.
Quickstart
pip install dbl-tempo
from tempo import TSDF
tsdf = TSDF(spark_df, ts_col="event_ts", partition_cols=["user_id"])
resampled = tsdf.resample(freq='min', func='mean')
Verify before relying
- Whether Tempo's operations scale efficiently to very large time series datasets or specific volume thresholds.
- Performance characteristics of AS OF joins compared to native Spark windowing on typical workloads.
- Support for custom aggregation functions beyond the documented 'floor', 'ceil', 'min', 'max', 'mean' options.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 333 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 4,221,234 / month, #2,360 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaProgramming Language :: PythonProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPy |
Evidence: dbl_tempo-0.1.30-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “spark time series operations”
- dbl-tempoTempo provides time series operations on Spark DataFrames, including…
- statsforecastStatsForecast provides fast implementations of statistical…
- mlforecastmlforecast trains machine learning models on time series data and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Information Analysis packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.
NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.
Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.
Install it if you use Snowflake and want to process data without moving it to your application layer.
See also dbldatagen · window-ops · pyspark-pandas · quinn · coreforecast · sparkaid · delta-spark · databricks-labs-remorph · databricks-test · dtw-python