impyla
Python client for the Impala distributed query engine
Decision gist · record as of 2026-08-14
Yes. impyla is a mature, actively maintained DB API 2.0 client for HiveServer2 with low install friction, permissive licensing, no known vulnerabilities, and broad authentication support. Install it if you need to query Impala or Hive from Python.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a running HiveServer2 instance accessible at the specified host and port.
- Optional Kerberos support requires system Kerberos libraries (libkrb5-dev on Ubuntu, krb5-devel on RHEL/CentOS).
- Low friction: pure Python wheel with four runtime dependencies (bitarray, thrift, thrift_sasl, pure-sasl).
License · maintenance · safety
Apache License, Version 2.0 (permissive) — Apache License, Version 2.0 (permissive) allows use, modification, and distribution with minimal restrictions; suitable for commercial and open-source projects.
last release 2026-06-19 (56 days) · last repo commit 2026-06-19 · 744 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 11,763,967 downloads/mo, #1,364 on PyPI
Alternatives
Verify before relying
pip install impyla
from impyla.dbapi import connect
conn = connect(host='my.host.com', port=21050)
cursor = conn.cursor()
cursor.execute('SELECT * FROM mytable LIMIT 100')
results = cursor.fetchall()- Performance characteristics and scalability limits for large result sets or high-concurrency scenarios.
- Compatibility with specific HiveServer2 versions or Impala releases beyond the general 'HiveServer2 compliant' claim.
- Whether pandas integration (as_pandas utility) is production-ready or primarily for exploratory use.
What it is and what it does
impyla provides a standard Python database interface to Impala and Hive query engines via HiveServer2. It implements PEP 249 (DB API 2.0), making it compatible with tools and libraries that expect standard Python database drivers. The package handles authentication (Kerberos, LDAP, SSL, JWT), result fetching, and schema introspection through a familiar cursor-based API.
Typical usage involves connecting to a HiveServer2 endpoint, executing SQL queries, and retrieving results as tuples or converting them to pandas DataFrames. It supports both synchronous iteration over result sets and bulk fetching, making it suitable for ad-hoc queries and data pipeline integration. The package has four runtime dependencies (bitarray, thrift, thrift_sasl, pure-sasl) and requires Python 3.8 or later.
Use it for
- Query Impala or Hive clusters from Python scripts for exploratory data analysis and ad-hoc SQL execution.
- Build ETL pipelines that extract data from distributed query engines and load into pandas for scikit-learn workflows.
- Integrate Impala/Hive queries into Python applications using standard DB API 2.0 patterns.
- Export query results to CSV by iterating over cursor rows and accessing column metadata.
- Connect securely to HiveServer2 with Kerberos or SSL authentication in enterprise environments.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
impyla is a mature, actively maintained DB API 2.0 client for HiveServer2 with low install friction, permissive licensing, no known vulnerabilities, and broad authentication support. Install it if you need to query Impala or Hive from Python.
Install
impyla on PyPI
Before you install
Low friction: pure Python wheel with four runtime dependencies (bitarray, thrift, thrift_sasl, pure-sasl). Active maintenance with a recent release 56 days ago and last commit 2026-06-19.
Requires a running HiveServer2 instance accessible at the specified host and port. Optional Kerberos support requires system Kerberos libraries (libkrb5-dev on Ubuntu, krb5-devel on RHEL/CentOS).
License in practice
Apache License, Version 2.0 (permissive) allows use, modification, and distribution with minimal restrictions; suitable for commercial and open-source projects.
Quickstart
pip install impyla
from impyla.dbapi import connect
conn = connect(host='my.host.com', port=21050)
cursor = conn.cursor()
cursor.execute('SELECT * FROM mytable LIMIT 100')
results = cursor.fetchall()
Verify before relying
- Performance characteristics and scalability limits for large result sets or high-concurrency scenarios.
- Compatibility with specific HiveServer2 versions or Impala releases beyond the general 'HiveServer2 compliant' claim.
- Whether pandas integration (as_pandas utility) is production-ready or primarily for exploratory use.
Package facts
| License | Apache License, Version 2.0 permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 4 packagesbitarraythriftthrift_saslpure-sasl |
| Maintenance | Actively maintained 56 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 11,763,967 / month, #1,364 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: impyla-0.24.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “hiveserver2 python client”
- impylaimpyla is a Python DB API 2.0-compliant client for querying…
- codewords-clientA Python client library for the Codewords API with built-in FastAPI…
- openapi-python-clientGenerates type-annotated Python HTTP clients from OpenAPI 3.0 and 3.1…
Give your agent the search over MCP, or paste the wish link into any chat.
More Database packages
psycopg2-binary is a PostgreSQL database adapter for Python that implements the DB API 2.0 specification, enabling Python applications to connect to and query PostgreSQL databases with thread-safe concurrent operations.
Python client library for connecting to and executing commands against Redis key-value stores, supporting both synchronous and asynchronous operations.
Install it if your application needs to interact with Redis; the only prerequisite is a running Redis server instance.
YDB Python SDK is the official client library for connecting to and querying YDB databases from Python applications.
Install it if you need to connect Python applications to YDB databases.
Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.
sqlparse tokenizes SQL text into a tree of statements, clauses, and expressions, and provides functions to split scripts, format queries, and inspect parsed tokens without validating dialect or syntax.
Install it if you need to manipulate, format, or analyze SQL text programmatically.
Provides base adapter protocols and shared functionality that database adapters use to integrate with dbt-core, handling connections, dialect translation, relation caching, and core interface management.
See also apache-airflow-providers-apache-impala · hdbcli · PyHive · pydynamodb · pyhdb · databricks-dbapi · PyAthena · trino · ibis-framework