{"enrichment":{"faq":[{"a":"ai-ml-data-science structures the complete workflow from problem framing through production deployment. Start with exploratory data analysis and feature engineering, establish baselines before complex models, then select and train models using tools like LightGBM or scikit-learn. The skill emphasizes leakage prevention, train-serve parity, and slice analysis during evaluation. Finally, document your model with contracts and lineage before deploying to production with MLOps pipelines.","q":"How do I build machine learning models from scratch with ai-ml-data-science?"},{"a":"ai-ml-data-science prioritizes leakage prevention as a core design principle. It covers feature engineering practices that maintain train-serve parity, ensuring features computed at training time match exactly what's available at inference. The skill guides you through proper data splitting, temporal validation for time series, and feature store setup with tools like Feast to enforce consistency across training and serving environments.","q":"What does ai-ml-data-science teach about preventing data leakage?"},{"a":"ai-ml-data-science covers LightGBM and CatBoost for tabular data, scikit-learn for classical methods, and Polars for efficient data transformation. For evaluation, it teaches slice analysis, baseline comparison, and metrics selection. The skill also covers experiment tracking with MLflow and Weights & Biases, hyperparameter tuning with Optuna and Ray, and model interpretability using SHAP and LIME.","q":"Which tools does ai-ml-data-science recommend for model selection and evaluation?"},{"a":"ai-ml-data-science structures MLOps around continuous training pipelines, drift monitoring with tools like Evidently, and reproducible experiment setup. It covers data contracts, lineage tracking, and model card documentation for operational handoff. The skill includes production deployment checklists and patterns for monitoring model performance and data drift in live systems.","q":"How does ai-ml-data-science address MLOps and production deployment?"},{"a":"ai-ml-data-science teaches SQL transformation workflows using SQLMesh for staging and feature computation. It emphasizes reproducible data pipelines that support both training and serving, with proper versioning and lineage tracking. The skill covers best practices for exploratory data analysis and feature engineering that maintain data quality and enable audit trails.","q":"What SQL and data transformation practices are included in ai-ml-data-science?"},{"a":"Yes. ai-ml-data-science includes time series forecasting with LightGBM, addressing temporal validation and drift concerns specific to sequential data. For classification, it covers handling class imbalance through proper evaluation metrics, baseline strategies, and slice analysis. Both topics are integrated into the broader lifecycle framework emphasizing reproducibility and production readiness.","q":"Does ai-ml-data-science cover time series forecasting and class imbalance?"}],"shadow_tags":["end-to-end-ml","tabular-modeling","feature-ops","production-readiness","data-quality-gates","model-governance","experiment-tracking","drift-detection","operational-handoff","baseline-first"],"summary_rewrite":"This skill structures the complete data science lifecycle\u2014from problem framing and exploratory analysis through feature pipelines and model evaluation to production deployment. It emphasizes baselines first, leakage prevention, train-serve parity, and reproducibility using tools like LightGBM, scikit-learn, and Polars. Covers SQL transformation with SQLMesh, experiment tracking, drift monitoring, and operational handoff patterns."},"files":[{"bytes":18436,"path":"frameworks/shared-skills/skills/ai-ml-data-science/SKILL.md","sha256":"b9d59e34b9500621f8adeadaf5a037c64ef4c4aeda729b164dbcc36a2c79896d","url":"https://skillfed.io/files/vasilyu1983/AI-Agents-public/ai-ml-data-science/79785576/SKILL.md"}],"id":"vasilyu1983/AI-Agents-public/ai-ml-data-science","links":{"html":"https://skillfed.io/vasilyu1983/AI-Agents-public/ai-ml-data-science","md":"https://skillfed.io/vasilyu1983/AI-Agents-public/ai-ml-data-science.md","repo":"https://github.com/vasilyu1983/AI-Agents-public"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":17,"language":"Python","last_updated":"2026-07-13","license":"MIT","name":"ai-ml-data-science","publisher":"vasilyu1983","stars":69},"relations":{"similar":[{"id":"vasilyu1983/AI-Agents-public/ai-mlops"},{"id":"vasilyu1983/AI-Agents-public/data-sql-optimization"},{"id":"vasilyu1983/AI-Agents-public/ai-ml-timeseries"},{"id":"vasilyu1983/AI-Agents-public/data-analytics-engineering"},{"id":"vasilyu1983/AI-Agents-public/data-lake-platform"},{"id":"wshobson/agents/dbt-transformation-patterns"},{"id":"ancoleman/ai-design-components/transforming-data"},{"id":"manutej/luxor-claude-marketplace/dbt-data-transformation"},{"id":"personamanagmentlayer/pcl/ai-architect-expert"},{"id":"personamanagmentlayer/pcl/dbt-expert"}]},"slug":{"owner":"vasilyu1983","repo":"AI-Agents-public","skill":"ai-ml-data-science"},"version":"79785576"}
