$npx skillfedfor your agent

ossdata

Scalable SWE datasets

With conditionsPyPI Artificial IntelligenceReleased Apr 2026234.4K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — ossdata-0.3.7-py3-none-any.whl
v0.3.7 · released 2026-04-20 · Python >=3.8 · 3 runtime deps: datasets, tqdm, alibabacloud-oss-v2

Yes, if you need scalable SWE datasets and are comfortable with the Hugging Face datasets ecosystem. Low install friction, permissive MIT license, active maintenance, and no known vulnerabilities make it a safe choice. The main uncertainty is whether the available datasets match your specific research or development needs—verify dataset availability before committing.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.8 or later; Alibaba OSS credentials may be needed depending on dataset source.
  • Low friction install with three straightforward dependencies.
  • Active maintenance as of 2026-04-20, released within the last 116 days.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; suitable for most projects.

last release 2026-04-20 (116 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 234,389 downloads/mo, #9,023 on PyPI

Verify before relying

pip install ossdata

from ossdata import load_dataset
dataset = load_dataset('dataset_name')
  • What specific SWE datasets are available and how they are organized.
  • Whether Alibaba OSS credentials are required for all datasets or only some.
  • Performance characteristics and scalability limits for large datasets.
  • Documentation and examples beyond the package summary.
Same gist for agents: .md · .json

What it is and what it does

OSSData is a Python package that wraps the Hugging Face datasets library to provide access to software engineering datasets, with backend support for Alibaba OSS cloud storage. It sits between your code and cloud-hosted SWE data, handling the mechanics of fetching and caching datasets locally.

The package depends on datasets for core dataset handling, tqdm for progress reporting, and alibabacloud-oss-v2 for cloud storage integration. It targets Python 3.8 and later and is actively maintained. The minimal description suggests it's designed for researchers and engineers who need standardized, reproducible access to SWE-focused datasets without managing storage infrastructure directly.

Use it for

  • Fetch and cache software engineering datasets for machine learning model training.
  • Access cloud-hosted SWE data without managing Alibaba OSS credentials directly.
  • Build reproducible research pipelines that depend on versioned, standardized datasets.
  • Integrate SWE datasets into data processing workflows using the Hugging Face ecosystem.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need scalable SWE datasets and are comfortable with the Hugging Face datasets ecosystem.

Low install friction, permissive MIT license, active maintenance, and no known vulnerabilities make it a safe choice. The main uncertainty is whether the available datasets match your specific research or development needs—verify dataset availability before committing.

Install

ossdata on PyPI

Before you install

Low friction install with three straightforward dependencies. Active maintenance as of 2026-04-20, released within the last 116 days.

Requires Python 3.8 or later; Alibaba OSS credentials may be needed depending on dataset source.

License in practice

MIT license permits commercial and private use with minimal restrictions; suitable for most projects.

Quickstart

pip install ossdata

from ossdata import load_dataset
dataset = load_dataset('dataset_name')

Verify before relying

  • What specific SWE datasets are available and how they are organized.
  • Whether Alibaba OSS credentials are required for all datasets or only some.
  • Performance characteristics and scalability limits for large datasets.
  • Documentation and examples beyond the package summary.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.8
Install frictionLow. Pure-Python wheel
Runtime dependencies
3 packages
datasetstqdmalibabacloud-oss-v2
MaintenanceActively maintained 116 days since the last release
First released
Downloads234,389 / month, #9,023 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3

Evidence: ossdata-0.3.7-py3-none-any.whl

Tags

Capabilities
software engineering datasetsscalable swe datacloud-backed datasetshugging face datasets integrationoss data loadingmachine learning dataset tools
Topics
datasetsmachine-learningsoftware-engineering

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “software engineering datasets”

  • ossdataProvides scalable datasets for software engineering research and…
  • swesmithSWE-smith generates large-scale software engineering training…
  • swebenchSWE-bench is a benchmark framework for evaluating language models on…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also datasets · kernels-data · pyspark-huggingface · alibabacloud-gateway-oss · alibabacloud-oss-v2 · oss2 · apache-airflow-providers-alibaba · ossfs · argilla