sagemaker-data-insights
Data Insights Library for Amazon SageMaker.
Decision gist · record as of 2026-08-14
No—not recommended for new projects. The package is abandoned (no updates since 2023-03-28) and licensed under a non-standard AWS Customer Agreement. While it has low install friction and no known vulnerabilities, the lack of maintenance means it will likely fall out of sync with SageMaker and its dependencies over time.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >=3.7; depends on numpy, pandas, scikit-learn, scipy, and sagemaker-scikit-learn-extension, all of which must be installed.
- Low install friction with a pure-Python wheel, but the package is abandoned—no updates since March 2023.
- Maintenance status is a significant concern for long-term reliability and compatibility with evolving SageMaker APIs.
License · maintenance · safety
AWS Customer Agreement (unclear) — Licensed under AWS Customer Agreement, which is not a standard open-source license. The unclear license treatment means you should review the agreement terms before using this package in production or redistributing it.
last release 2023-03-28 (1235 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 579,163 downloads/mo, #5,919 on PyPI
Alternatives
Verify before relying
pip install sagemaker-data-insights
import sagemaker_data_insights
# Use with SageMaker datasets to generate statistical summaries- What specific ML statistics does the package compute (e.g., distribution analysis, missing-value patterns, feature correlations)?
- Does it work with current SageMaker SDK versions, or has the abandoned status caused API drift?
- Are there known compatibility issues with recent versions of numpy, pandas, or scikit-learn?
What it is and what it does
sagemaker-data-insights is a SageMaker-integrated library that analyzes datasets and produces statistics relevant to machine learning workflows. It sits on top of standard data science libraries (numpy, pandas, scikit-learn, scipy) and the sagemaker-scikit-learn-extension to compute summaries of data characteristics.
The package is designed to help ML practitioners understand their datasets before and during model development—surfacing patterns, anomalies, or quality issues that might affect training. However, it has been abandoned since its latest release on 2023-03-28 and receives no maintenance, which means it may not work reliably with newer versions of its dependencies or with current SageMaker APIs.
Use it for
- Profile raw datasets in SageMaker to identify data quality issues before training a model.
- Generate statistical summaries of features to inform feature engineering decisions.
- Diagnose dataset imbalances or missing-value patterns that could affect model performance.
- Integrate data insights into SageMaker notebook workflows for exploratory data analysis.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
No—not recommended for new projects.
The package is abandoned (no updates since 2023-03-28) and licensed under a non-standard AWS Customer Agreement. While it has low install friction and no known vulnerabilities, the lack of maintenance means it will likely fall out of sync with SageMaker and its dependencies over time.
Install
sagemaker-data-insights on PyPI
Before you install
Low install friction with a pure-Python wheel, but the package is abandoned—no updates since March 2023. Maintenance status is a significant concern for long-term reliability and compatibility with evolving SageMaker APIs.
Requires Python >=3.7; depends on numpy, pandas, scikit-learn, scipy, and sagemaker-scikit-learn-extension, all of which must be installed.
License in practice
Licensed under AWS Customer Agreement, which is not a standard open-source license. The unclear license treatment means you should review the agreement terms before using this package in production or redistributing it.
Quickstart
pip install sagemaker-data-insights
import sagemaker_data_insights
# Use with SageMaker datasets to generate statistical summaries
Verify before relying
- What specific ML statistics does the package compute (e.g., distribution analysis, missing-value patterns, feature correlations)?
- Does it work with current SageMaker SDK versions, or has the abandoned status caused API drift?
- Are there known compatibility issues with recent versions of numpy, pandas, or scikit-learn?
Package facts
| License | AWS Customer Agreement unclear |
| Python support | Supports the current Python release >=3.7 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagesnumpypsutilpandasscikit-learnscipysagemaker-scikit-learn-extension |
| Maintenance | Abandoned 1,235 days since the last release |
| First released | |
| Downloads | 579,163 / month, #5,919 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - Beta |
Evidence: sagemaker_data_insights-0.4.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “sagemaker data statistics”
- sagemaker-data-insightsComputes ML-relevant statistical summaries of datasets, integrating…
- sagemaker-datawranglerProvides a Python library for Amazon SageMaker Data Wrangler,…
- sagemaker-studioA Python SDK for accessing Amazon SageMaker Unified Studio…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also sagemaker-serve · whylogs · sagemaker · sagemaker-train · sagemaker-schema-inference-artifacts · sagemaker-training · sagemaker-datawrangler · sagemaker-inference · sagemaker-feature-store-pyspark-3.1 · sagemaker-scikit-learn-extension