$npx skillfedfor your agent

datasieve

This package implements a flexible data pipeline to help organize row removal (e.g. outlier removal) and feature modification (e.g. PCA)

With conditionsPyPI Scientific/EngineeringReleased May 2025214.1K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — datasieve-0.1.9-py3-none-any.whl
v0.1.9 · released 2025-05-11 · Python <4.0,>=3.8.1 · 2 runtime deps: pandas, scikit-learn

Yes, if you need to coordinate row or feature removals across X, y, and sample_weight in a single pipeline. The low install friction, permissive license, and absence of known vulnerabilities make it safe to try. However, the aging maintenance status (460 days since last release) means you should verify compatibility with your specific pandas and scikit-learn versions before relying on it in production.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires pandas and scikit-learn as runtime dependencies; Python 3.8.1 or later.
  • Low install friction with a pure-Python wheel.
  • Maintenance is aging (460 days since last release), but the package remains compatible with current Python versions (3.8.1 through 3.13) and carries no known vulnerabilities.

License · maintenance · safety

MIT (permissive) — MIT license is permissive; you can use, modify, and distribute DataSieve freely in commercial or private projects with minimal restrictions.

last release 2025-05-11 (460 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 214,114 downloads/mo, #9,424 on PyPI

Verify before relying

from datasieve.pipeline import Pipeline
import datasieve.transforms as dst

feature_pipeline = Pipeline([
    ("detect_constants", dst.VarianceThreshold(threshold=0)),
    ("svm", dst.SVMOutlierExtractor())
])

X, y, sample_weight = feature_pipeline.fit_transform(X, y, sample_weight)
  • Whether the package is actively maintained or in stable maintenance mode despite the aging status.
  • Real-world performance characteristics when handling large datasets with complex pipelines.
  • Compatibility guarantees with recent pandas and scikit-learn versions beyond what classifiers declare.
Same gist for agents: .md · .json

What it is and what it does

DataSieve is a Pipeline extension that coordinates transformations across feature arrays (X), target arrays (y), and sample weights simultaneously. Unlike standard pipelines that only transform X, DataSieve propagates row removals and feature modifications through all three arrays in lockstep—so when an outlier detection step removes rows from X, it automatically removes the corresponding rows from y and sample_weight.

The package includes built-in transforms for common tasks like variance-based feature filtering, SVM-based outlier detection with automatic removal, and PCA with feature renaming. It supports wrapping transforms directly, and allows custom transforms that manipulate any combination of X, y, and sample_weight. A key feature is outlier flagging without removal: you can fit a pipeline and call transform with outlier_check=True to get a binary vector marking outliers while keeping all data intact.

Use it for

  • Remove outliers from training data while keeping y and sample_weight synchronized for supervised learning.
  • Apply feature selection or dimensionality reduction and automatically track feature name changes.
  • Build multi-stage preprocessing pipelines that filter rows based on complex criteria across multiple arrays.
  • Flag anomalies in new data without removing them, returning both transformed features and outlier indicators.
  • Wrap transforms in a coordinated pipeline that handles weighted or imbalanced datasets with row-level filtering.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need to coordinate row or feature removals across X, y, and sample_weight in a single pipeline.

The low install friction, permissive license, and absence of known vulnerabilities make it safe to try. However, the aging maintenance status (460 days since last release) means you should verify compatibility with your specific pandas and scikit-learn versions before relying on it in production.

Install

datasieve on PyPI

Before you install

Low install friction with a pure-Python wheel. Maintenance is aging (460 days since last release), but the package remains compatible with current Python versions (3.8.1 through 3.13) and carries no known vulnerabilities.

Requires pandas and scikit-learn as runtime dependencies; Python 3.8.1 or later.

License in practice

MIT license is permissive; you can use, modify, and distribute DataSieve freely in commercial or private projects with minimal restrictions.

Quickstart

from datasieve.pipeline import Pipeline
import datasieve.transforms as dst

feature_pipeline = Pipeline([
    ("detect_constants", dst.VarianceThreshold(threshold=0)),
    ("svm", dst.SVMOutlierExtractor())
])

X, y, sample_weight = feature_pipeline.fit_transform(X, y, sample_weight)

Verify before relying

  • Whether the package is actively maintained or in stable maintenance mode despite the aging status.
  • Real-world performance characteristics when handling large datasets with complex pipelines.
  • Compatibility guarantees with recent pandas and scikit-learn versions beyond what classifiers declare.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release <4.0,>=3.8.1
Install frictionLow. Pure-Python wheel
Runtime dependencies
2 packages
pandasscikit-learn
MaintenanceAging 460 days since the last release
First released
Downloads214,114 / month, #9,424 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.9

Evidence: datasieve-0.1.9-py3-none-any.whl

Tags

Capabilities
pipeline with y and sample_weightoutlier removal pipelinefeature engineering pipelinedata preprocessing with row filteringcoordinated feature and sample transformationdimensionality reduction pipelinedata transformation pipeline
Topics
data-preprocessingoutlier-detectionpipeline-extension

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “pipeline with y and sample_weight”

  • datasieveDataSieve extends scikit-learn's Pipeline to handle row and feature…
  • kfp-pipeline-specProvides the pipeline specification and protobuf definitions for…
  • kedro-vizKedro-Viz is an interactive web-based visualization tool for Kedro…

Give your agent the search over MCP, or paste the wish link into any chat.

More Scientific/Engineering packages

numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
pandas Worth it
PyPI · Scientific/Engineering · released Jul 2026

pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.

BSD-3-Clausecompiled wheel · 3.11+
769.1Mdownloads / mo
scipy Worth it
PyPI · Libraries · released Jun 2026

scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.

BSD-3-Clausecompiled wheel · 3.12+
449.0Mdownloads / mo
scikit-learn Worth it
PyPI · Software Development · released Jun 2026

scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.

Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.

BSD-3-Clausecompiled wheel · 3.11+
235.5Mdownloads / mo
dill Worth it
PyPI · Software Development · released Jan 2026

dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.

BSD-3-Clausepure Python · 3.9+
208.1Mdownloads / mo
multiprocess Worth it
PyPI · Software Development · released Jan 2026

Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.

Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.

BSD-3-Clausepure Python · 3.9+
202.7Mdownloads / mo

See also feature-engine · sklearn-pandas · sklearndf · category-encoders · hampel · spark-sklearn · umap-learn · sagemaker-scikit-learn-extension · azureml-dataprep