~ / skills / data-science

data-science/

435 total · showing first 300
productivity 3,805software-engineering 2,099writing 1,172ai-agents 1,111security 946healthcare 712finance 467data-science 435marketing 344hr 334devops 289research 253product-design 193media 181legal 143education 89sales 76project-management 64ecommerce 48support 40manufacturing 38gaming 17

jupyter-live-kernel

Iterative Python via live Jupyter kernel (hamelnb).

Unverified214,858

nemo-curator

GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic…

Unverified214,858

pdf

Use this skill whenever the user wants to do anything with PDF files. This includes reading or extracting text/tables from PDFs, combining or merging multiple PDFs into one…

Link only161,058

pdf

Comprehensive PDF manipulation toolkit for extracting text and tables, creating new PDFs, merging/splitting documents, and handling forms. When Claude needs to fill in a PDF form…

Unverified67,708

googlebigquery-automation

Automate Google BigQuery tasks via Rube MCP (Composio): run SQL queries, explore datasets and metadata, execute MBQL queries via Metabase integration. Always search tools first…

Unverified67,708

azure-ai-ml-py

Azure Machine Learning SDK v2 for Python. Use for ML workspaces, jobs, models, datasets, compute, and pipelines.

Unverified43,234

azure-storage-file-datalake-py

Azure Data Lake Storage Gen2 SDK for Python. Use for hierarchical file systems, big data analytics, and file/directory operations.

Unverified43,234

azure-monitor-query-py

Azure Monitor Query SDK for Python. Use for querying Log Analytics workspaces and Azure Monitor metrics.

Unverified43,234

llm-prompt-optimizer

Use when improving prompts for any LLM. Applies proven prompt engineering techniques to boost output quality, reduce hallucinations, and cut token usage.

Unverified43,234

local-llm-expert

Master local LLM inference, model selection, VRAM optimization, and local deployment using Ollama, llama.cpp, vLLM, and LM Studio. Expert in quantization formats (GGUF, EXL2) and…

Unverified43,234

python-pptx-generator

Generate complete Python scripts that build polished PowerPoint decks with python-pptx and real slide content.

Unverified43,228

spreadsheets

Create, edit, analyze, clean, or convert spreadsheets including XLSX, CSV, TSV, formulas, charts, and tabular reports.

Unverified39,782

uv-package-manager

Master the uv package manager for fast Python dependency management, virtual environments, and modern Python project workflows. Use when setting up Python projects, managing…

Unverified37,902

airflow-dag-patterns

Build production Apache Airflow DAGs with best practices for operators, sensors, testing, and deployment. Use when creating data pipelines, orchestrating workflows, or scheduling…

Unverified37,902

azure-pricing

Fetches real-time Azure retail pricing using the Azure Retail Prices API (prices.azure.com) and estimates Copilot Studio agent credit consumption. Use when the user asks about…

Unverified36,563

dataverse-python-advanced-patterns

Generate production code for Dataverse SDK using advanced patterns, error handling, and optimization techniques.

Unverified36,563

medchem

Medicinal chemistry filters for compound triage. Apply drug-likeness rules (Lipinski, Veber, CNS), structural alert catalogs (PAINS, NIBR, ChEMBL), complexity metrics, and the…

Unverified30,891

geopandas

Python library for working with geospatial vector data including shapefiles, GeoJSON, and GeoPackage files. Use when working with geographic data for spatial analysis, geometric…

Unverified30,891

polars

High-performance DataFrame library for Python ETL, analytics, and pandas migration. Use for expression-based data manipulation with lazy query optimization, parallel execution…

Unverified30,891

scikit-learn

Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality…

Unverified30,891

aeon

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity…

Unverified30,891

anndata

Data structure for annotated matrices in single-cell analysis. Use when working with .h5ad files or integrating with the scverse ecosystem. This is the data format skill—for…

Unverified30,891

arboreto

Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell…

Unverified30,891

cellxgene-census

Query the CZ CELLxGENE Census programmatically for versioned public single-cell and spatial transcriptomics data. Use when you need population-scale cell metadata, gene…

Unverified30,891

dask

Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file…

Unverified30,891

deepchem

Molecular ML with diverse featurizers and pre-built datasets. Use for property prediction (ADMET, toxicity) with traditional ML or GNNs when you want extensive featurization…

Unverified30,891

depmap

Query the Cancer Dependency Map (DepMap) for cancer cell line gene dependency scores (CRISPR Chronos), drug sensitivity data, and gene effect profiles. Use for identifying…

Unverified30,891

geniml

This skill should be used when working with genomic interval data (BED files) for machine learning tasks. Use for training region embeddings (Region2Vec, BEDspace), single-cell…

Unverified30,891

gtars

High-performance toolkit for genomic interval analysis in Rust with Python bindings. Use when working with genomic regions, BED files, coverage tracks, overlap detection…

Unverified30,891

molfeat

Molecular featurization for ML (100+ featurizers). ECFP, MACCS, descriptors, pretrained models (ChemBERTa), convert SMILES to features, for QSAR and molecular ML.

Unverified30,891

pathml

Full-featured computational pathology toolkit. Use for advanced WSI analysis including multiplexed immunofluorescence (CODEX, Vectra), nucleus segmentation, tissue graph…

Unverified30,891

seaborn

Statistical visualization with pandas integration. Use for quick exploration of distributions, relationships, and categorical comparisons with attractive defaults. Best for box…

Unverified30,891

shap

Model interpretability and explainability using SHAP (SHapley Additive exPlanations). Use this skill when explaining machine learning model predictions, computing feature…

Unverified30,891

statsmodels

Statistical models library for Python. Use when you need specific model classes (OLS, GLM, mixed models, ARIMA) with detailed diagnostics, residuals, and inference. Best for…

Unverified30,891

tiledbvcf

Efficient storage and retrieval of genomic variant data using TileDB. Scalable VCF/BCF ingestion, incremental sample addition, compressed storage, parallel queries, and export…

Unverified30,891

experimental-design

Design experiments and studies BEFORE data is collected — choosing a design, randomizing, blocking, and laying out treatment combinations so the results will actually be…

Unverified30,891

geomaster

Comprehensive geospatial science skill covering remote sensing, GIS, spatial analysis, machine learning for earth observation, and 30+ scientific domains. Supports satellite…

Unverified30,891

pennylane

Hardware-agnostic quantum ML framework with automatic differentiation. Use when training quantum circuits via gradients, building hybrid quantum-classical models, or needing…

Unverified30,891

polars-bio

High-performance genomic interval operations and bioinformatics file I/O on Polars DataFrames. Overlap, nearest, merge, coverage, complement, subtract for BED/VCF/BAM/GFF…

Unverified30,891

statistical-analysis

Guided statistical analysis for research data - test selection, assumption checking, effect sizes, power analysis, Bayesian alternatives, and APA-formatted reporting. Use…

Unverified30,891

pdf

Read, extract (text/tables), create, merge/split/rotate, watermark, encrypt,

Unverified25,903

xlsx

Read, create, or edit Excel spreadsheets (.xlsx/.xlsm) — sheet data,

Unverified25,903

cli-creator

Build a composable CLI for Codex from API docs, an OpenAPI spec, existing curl examples, an SDK, a web app, an admin tool, or a local script. Use when the user wants Codex to…

Unverified23,696

cohort-analysis

Perform cohort analysis on user engagement data — retention curves, feature adoption trends, and segment-level insights. Use when analyzing user retention by cohort, studying…

Unverified23,669

candlestick

Candlestick pattern recognition engine, pure pandas vectorized implementation of 15 classic candlestick patterns (5 single-candle + 5 double-candle + 4 triple-candle + 1 trend…

Unverified22,765

chanlun

基于缠论(缠中说禅)的形态识别引擎,使用czsc库自动检测K线分型、笔、中枢,并生成一买/一卖/二买/二卖/三买/三卖等买卖点信号。支持多周期分析和形态分类(3/5/7/9/11笔形态)。

Unverified22,765

correlation-analysis

Correlation and cointegration analysis — co-movement discovery, deep return-correlation analysis, sector clustering, realized correlation, Engle-Granger / Johansen cointegration…

Unverified22,765

elliott-wave

Elliott Wave Theory signal engine. Detects swing points through Zigzag, matches 5-wave impulse and 3-wave corrective structures, validates them with Fibonacci wave relationships…

Unverified22,765

event-driven

Event-driven strategy based on sentiment-scored signals from news, announcements, and macro events. The LLM acts as the NLP engine, and event data follows a CSV schema.

Unverified22,765

ichimoku

Ichimoku Kinko Hyo five-line system signal engine. A standalone Japanese technical-analysis school that generates trading signals from Tenkan/Kijun crossovers, cloud position…

Unverified22,765

minute-analysis

Minute-level data analysis and backtesting. Retrieves minute candlesticks through OKX/Tushare/yfinance and can be used both for analysis and as input to the backtest engine.

Unverified22,765

ml-strategy

Machine-learning predictive strategy based on sklearn walk-forward training, feature engineering, and signal generation. Suitable for any OHLCV data.

Unverified22,765

multi-factor

Multi-factor cross-sectional stock ranking. Combines factor standardization, equal-weight or IC-weighted scoring, and TopN portfolio construction. Suitable for multi-instrument…

Unverified22,765

okx-market

OKX cryptocurrency market data interface. Uses the OKX V5 REST API to retrieve spot, derivatives, index, and other crypto market data, including real-time prices, candlesticks…

Unverified22,765

pair-trading

Pair trading strategy. Trades mean reversion using the spread/ratio Z-score of two correlated instruments. Requires at least two instruments.

Unverified22,765

risk-analysis

Risk measurement and stress testing — VaR/CVaR/max drawdown calculation, Monte Carlo simulation, extreme-value tail-risk analysis, and historical scenario stress testing.

Unverified22,765

smc

Smart Money Concepts (ICT) signal engine. Uses the smartmoneyconcepts library to implement institutional-trading-school analysis of BOS, ChoCH, FVG, and order blocks (OB).

Unverified22,765

technical-basic

Core technical indicator collection (trend EMA/ADX + mean-reversion BB/RSI + volume-price OBV/volume ratio), generates a composite signal via three-dimensional voting. Pure…

Unverified22,765

tushare

tushare是一个财经数据接口包,拥有丰富的数据内容,如股票、基金、期货、数字货币等行情数据,公司财务、基金经理等基本面数据。该模块通过标准化API方式统一了数据资产的对外服务方式,以帮助有需要的技术用户更实时、简洁、轻量的使用相关数据。

Unverified22,765

volatility

Volatility strategy. Trades mean reversion based on percentile ranking of historical volatility (HV). Suitable for any OHLCV data.

Unverified22,765

senior-data-scientist

World-class senior data scientist skill specialising in statistical modeling, experiment design, causal inference, and predictive analytics. Covers A/B testing (sample sizing…

Unverified22,559

statistical-analyst

Run hypothesis tests, analyze A/B experiment results, calculate sample sizes, and interpret statistical significance with effect sizes. Use when you need to validate whether…

Unverified22,559

strict-api

Use when the user says 'no hallucinations', 'verify APIs', 'reality check', or 'don't invent functions'. Prevents the agent from calling methods, imports, or variables that do…

Unverified22,559

universal-scraping-architect

Use for web scraping, crawling, document extraction, API parsing, or building validation-heavy data pipelines using Firecrawl or local Python scripts.

Unverified22,559

scrape

This skill searches job portals using the installed portal-search CLIs in

Unverified22,441

timescaledb

TimescaleDB - PostgreSQL extension for high-performance time-series and event data analytics, hypertables, continuous aggregates, compression, and real-time analytics

Unverified22,132

bigquery-ai-ml

This skill defines the usage and rules for BigQuery AI/ML functions,

Unverified20,613

data-analyst

Data analysis expert for statistics, visualization, pandas, and exploration

Unverified18,014

ml-engineer

Machine learning engineer expert for PyTorch, scikit-learn, model evaluation, and MLOps

Unverified18,014

bigquery-ai-ml

BigQuery integrates with Vertex AI to provide powerful machine learning and

Unverified14,785

agent-platform-eval-flywheel

Help users evaluate and iteratively improve GenAI models and agents using

Unverified14,785

agent-platform-model-registry

This skill provides instructions for managing machine learning models in the

Unverified14,785

bigquery-basics

BigQuery is a serverless, AI-ready data platform that enables high-speed

Unverified14,785

bigquery-bigframes

Avoid .topandas(): You MUST NOT use .topandas() to download the

Unverified14,785

google-cloud-solution-agentic-ai-data-science-workflow

This skill guides agents through the workflow to design and implement a

Unverified14,785

golden-jupyter

Use when testing the golden_jupyter golden build

Unverified14,454

golden-jupyter-dir

Use when testing the golden_jupyter_dir golden build

Unverified14,454

golden-jupyter-kw

Use when testing the golden_jupyter_kw golden build

Unverified14,454

golden-jupyter-topics

Use when testing the golden_jupyter_topics golden build

Unverified14,454

fba-simulator

The fba-simulator skill executes constraint-based metabolic simulations on

Unverified13,799

flux-analyzer

The flux-analyzer skill transforms raw FBA output into actionable biological

Unverified13,799

quantum-qiskit

Reference qiskit 2.x patterns for variational quantum machine learning. Covers data-encoding feature maps, variational quantum classifier (VQC) training, variational quantum…

Unverified13,799

auto-review-loop-minimax

Autonomous multi-round research review loop using MiniMax API. Use when you want to use MiniMax instead of Codex MCP for external review. Trigger with \"auto review loop…

Unverified13,392

maf-online-endpoint

Deploy a Microsoft Agent Framework (MAF) workflow as a managed online endpoint to an Azure ML workspace or an Azure AI Foundry hub-based project. Wraps any workflow into an…

Unverified11,183

transformers-js

Use Transformers.js to run state-of-the-art machine learning models directly in JavaScript/TypeScript. Supports NLP (text classification, translation, summarization), computer…

Unverified10,814

evaluating-code-models

Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing…

Unverified10,718

creative-thinking-for-research

Applies cognitive science frameworks for creative thinking to CS and AI research ideation. Use when seeking genuinely novel research directions by leveraging combinatorial…

Unverified10,718

langsmith-observability

LLM observability platform for tracing, evaluation, and monitoring. Use when debugging LLM applications, evaluating model outputs against datasets, monitoring production systems…

Unverified10,718

prompt-guard

Meta's 86M prompt injection and jailbreak detector. Filters malicious prompts and third-party data for LLM apps. 99%+ TPR, <1% FPR. Fast (<2ms GPU). Multilingual (8 languages).…

Unverified10,718

ray-data

Scalable data processing for ML workloads. Streaming execution across CPU/GPU, supports Parquet/CSV/JSON/images. Integrates with Ray Train, PyTorch, TensorFlow. Scales from…

Unverified10,718

ray-train

Distributed training orchestration across clusters. Scales PyTorch/TensorFlow/HuggingFace from laptop to 1000s of nodes. Built-in hyperparameter tuning with Ray Tune, fault…

Unverified10,718

sentence-transformers

Framework for state-of-the-art sentence, text, and image embeddings. Provides 5000+ pre-trained models for semantic similarity, clustering, and retrieval. Supports multilingual…

Unverified10,718

ml-pipeline

Designs and implements production-grade ML pipeline infrastructure: configures experiment tracking with MLflow or Weights & Biases, creates Kubeflow or Airflow DAGs for training…

Unverified10,585

pandas-pro

Performs pandas DataFrame operations for data analysis, manipulation, and transformation. Use when working with pandas DataFrames, data cleaning, aggregation, merging, or time…

Unverified10,585

kb-retriever

面向本地知识库目录的检索和问答助手。核心流程:(1)分层索引导航 (2)遇到PDF/Excel时必须先读取references学习处理方法 (3)处理文件后再检索。按文件类型组合使用 grep、Read、pdfplumber、pandas 进行渐进式检索,避免整文件加载。用户问题涉及"从知识库目录回答问题/检索信息/查资料"时使用。

Unverified9,505

code-exec-fallback

Fallback pattern for executing Python code when execute_code_sandbox fails

Unverified6,736

code-exec-fallback-266cba

Fallback workflow for reliable code execution when sandbox fails repeatedly

Unverified6,736

code-execution-fallback

Handle code execution failures with fallback strategies and anchored workspace paths

Unverified6,736

code-execution-fallback-e81068

Fallback workflow for executing Python code when execute_code_sandbox fails repeatedly

Unverified6,736

execute-code-fallback

Fallback from execute_code_sandbox to run_shell when e2b sandbox is unavailable

Unverified6,736

fallback-python-execution

Reliable Python execution workflow when execute_code_sandbox or shell_agent fail

Unverified6,736

fallback-python-shell

Use run_shell with Python heredoc when execute_code_sandbox or read_file fail

Unverified6,736

fallback-script-execution

Two-step script execution workflow for debugging when shell_agent and execute_code_sandbox consistently fail

Unverified6,736

prioritize-context-data

Ensures agents check and use provided context files for data before attempting external searches.

Unverified6,736

python-debug-execution-911f17

Debug Python script execution failures by capturing full tracebacks and verifying working directory

Unverified6,736

python-heredoc-fallback

Use shell Python heredoc as fallback when code sandbox fails

Unverified6,736

python-shell-workaround

Workaround for executing Python code with external packages when sandbox fails

Unverified6,736

python-spreadsheet-debug

Systematic Python debugging workflow for spreadsheet tasks to isolate environment issues from script logic

Unverified6,736

resilient-spreadsheet-workflow

Resilient multi-step workflow for spreadsheet processing when execute_code_sandbox fails, using shell_agent exploration, file-based scripts, and verification steps

Unverified6,736

sandbox-exec-fallback

Fallback pattern for executing Python code when execute_code_sandbox fails by writing to file and running via shell

Unverified6,736

sandbox-execution-fallback

Recover from execute_code_sandbox failures by writing Python scripts to files and executing via run_shell

Unverified6,736

sandbox-failure-recovery

Recover from execute_code_sandbox failures by writing code to file and executing via run_shell

Unverified6,736

sandbox-fallback-execution-048d5a

Fallback method to execute Python code when execute_code_sandbox fails with e2b errors

Unverified6,736

sandbox-fallback-python-e5d7ae

Fallback to run_shell with Python heredoc when execute_code_sandbox fails

Unverified6,736

shell-python-heredoc

Execute complex Python code via run_shell heredoc when execute_code_sandbox fails

Unverified6,736

web-search

This skill should be used when users need to search the web for information, find current content, look up news articles, search for images, or find videos. It uses DuckDuckGo's…

Unverified6,032

n8n-code-python

Write Python code in n8n Code nodes. Use when writing Python in n8n, using _input/_json/_node syntax, working with standard library, or need to understand Python limitations in…

Unverified5,789

large-file-parquet-analysis-and-highlight

当Excel文件总行数超过1万行时,通过转换为Parquet格式提升读取性能,提取目标指标并计算最大值,最后将结果输出为Excel并对特定行进行高亮标注。

Unverified4,710

category-filtering-and-difficulty-analysis

对Excel数据进行自定义分类统计、交叉分析与可视化,并基于多维度指标(如文本长度、术语密度、正则匹配等)进行综合评分与分级,适用于多类别数据分布统计及文本内容难度/质量评估场景。

Unverified4,710

category-statistics

提取指定类别列并统计各类别数量与占比,生成高分辨率的柱状图、饼图等组合可视化报告,适用于分类数据的分布情况分析。

Unverified4,710

categorical-comparison-analysis

对两类分类数据进行对比分析,统计数量差异与比例关系并生成可视化图表。

Unverified4,710

numeric-extraction-and-distribution-analysis

从带单位的字符串列中提取数值并清洗,生成包含直方图、饼图、条形图和累积分布图的多维度综合分布可视化图表,用于展示数据的集中趋势与分布特征。

Unverified4,710

excel-multi-sheet-threshold-analysis

统计多Sheet Excel总行数并根据规模选择处理策略,提取特定维度信息进行去重统计,并生成摘要与明细报表。

Unverified4,710

formatted-export-with-parquet

从多Sheet Excel文件中识别指定条件的记录,并将筛选结果以整行标红格式导出为Excel文件,适用于数据清洗、条件筛选与可视化标记场景。

Unverified4,710

grouped-statistics

对多 Sheet 的 Excel 文件进行行数统计、数据合并与前向填充。

Unverified4,710

statistical-distribution-and-outlier-analysis

执行数值型数据的分布分析与异常值检测,支持通过正则表达式从文本中提取误差项并生成高分辨率的箱线图与直方图报告。

Unverified4,710

invalid-data-cleaning

用于大规模Excel数据的预处理,通过统计总行数判断是否转换为Parquet格式以提升读写效率,并使用正则表达式清洗指定文本列(如仅保留中文字符),最后导出清洗后的文件并提供下载链接。

Unverified4,710

large-excel-analysis-and-formatting

用于处理多Sheet大型Excel文件,支持大文件Parquet格式转换提速,并使用openpyxl生成带条件高亮和自定义样式的格式化Excel报告及下载链接。

Unverified4,710

line-chart-visualization

提取结构化数据并进行特征清洗与聚类分析,生成包含趋势对比、分布特征与参数敏感性的多维度综合可视化图表,适用于各类趋势预测与多维对比场景。

Unverified4,710

multi-file-excel-parquet-analysis

读取多 Sheet Excel 文件并统计规模,支持大文件向 Parquet 格式转换、分类数据统计及可视化报告生成。

Unverified4,710

multi-sheet-reading-and-analysis

用于读取多工作表Excel文件,动态评估数据量以启用Parquet大文件优化,并执行正则清洗、分类汇总、线性拟合及生成带格式的图表与结果文件。

Unverified4,710

excel-outlier-detection-and-highlighting

识别 Excel 中的超限数值与错误单元格并进行高亮标注。

Unverified4,710

outlier-detection-and-quality-assessment

执行全面的异常值检测与数据质量评估,利用 IQR 方法识别异常值并结合偏度、峰度分析数据分布特征,适用于非正态分布数据的预处理阶段。

Unverified4,710

pdf-analysis

PDF 文档解析。自动区分文字型 PDF 与扫描型 PDF,覆盖:文本/表格提取、多页全量扫描、嵌入图表 caption、单位感知数值计算。

Unverified4,710

pie-chart-data-analysis

对多Sheet Excel或CSV数据进行分类汇总统计,自动识别关键字段并生成包含占比、数值及美化饼图的可下载分析报告。

Unverified4,710

pivot-table-cross-analysis

利用交叉表与热力图对分类数据进行多维度占比分析,适用于奖项分布、绩效评估或市场占有率等结构化数据的清洗与可视化。

Unverified4,710

ppt-analysis

PPT (.pptx/.ppt) 全量解析。覆盖:所有 slide 文本/表格/图表提取、嵌入图片 caption、纯图片 slide 渲染识别、数据标签提取。

Unverified4,710

excel-conditional-filtering-optimization

根据多维数值条件筛选 Excel 数据并导出结果,支持大规模数据的自动性能优化处理。

Unverified4,710

single-sheet-reading-and-analysis

读取并解析单个Excel工作表数据,支持合并单元格处理、数据清洗、交叉分析及多维度可视化,适用于需要从单表中提取关键指标并进行趋势模拟与图表生成的场景。

Unverified4,710

sn-da-excel-workflow

Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV…

Unverified4,710

sn-da-large-file-analysis

万行以上 Excel 数据集的高性能分析引擎。提供 openpyxl read_only 流式读取(iter_rows 支持 10 万行以上)、Parquet 转换加速、内存优化、分块处理和大文件写入模式。**遇到以下任一情况就主动使用本 skill**:①数据行数 ≥ 10k(由 sn-da-excel-workflow…

Unverified4,710

sn-search-code

用于查找代码示例、开源项目、GitHub Issue、技术问答、开发者讨论、HuggingFace 模型/数据集/Space。

Unverified4,710

sn-search-social-en

用于搜索英文社交平台,包括 Reddit 帖子、Twitter/X 推文和 YouTube 视频。

Unverified4,710

excel-multi-sheet-dynamic-analysis

用于分析包含多个Sheet的Excel文件,动态判断数据量级以决定是否转换为Parquet进行大文件处理,并支持跨Sheet的特定字段统计、数据清洗、交叉分析与可视化,最终生成带下载链接的汇总报告。

Unverified4,710

stacked-chart-visualization

处理包含百分比字符串的分类占比数据,通过补全缺失维度并生成堆叠柱状图,直观展示多维度构成随时间或分类的变化趋势。

Unverified4,710

large-file-conditional-formatting

根据Excel总行数自动切换Parquet加速读取,计算特定维度的时间序列平均值,并使用openpyxl输出带有条件格式(如低于均值标绿)和自定义样式的分析报告。

Unverified4,710

excel-threshold-analysis-and-styling

根据 Excel 数据量级自动判断处理策略,执行数值列清洗、条件过滤,并使用 openpyxl 对符合条件的单元格进行样式标记与导出。

Unverified4,710

time-series-and-categorical-analysis

对时间序列或分类数据进行多维度趋势分析、百分比清洗、绩效分级建模与预测,并生成高分辨率的可视化综合报告,适用于业务指标监控与预测场景。

Unverified4,710

trend-analysis

基于多维度数据进行分级评估与趋势预测,通过设定差异化增长率计算预测值,并生成对比可视化图表,适用于绩效评估、目标设定等场景。

Unverified4,710

word-analysis

Word (.docx/.doc) 文档全量解析。覆盖:正文/段落文本提取、表格数据提取、高亮/颜色格式读取、多文件汇总对比、嵌入图片转 caption。

Unverified4,710

technology-selection

Guides technology selection and implementation of AI and ML features in .NET 8+ applications using ML.NET, Microsoft.Extensions.AI (MEAI), Microsoft Agent Framework (MAF), GitHub…

Unverified4,615

chart-rendering

Visualise the result of an analysis as a chart (line, bar, area, scatter, etc.). Use when the user asks to "plot...", "chart...", "show me the trend of...", "visualise...", or…

Unverified4,463

python-repl

Interactive Python REPL automation with common helpers and best practices

Unverified4,354

stock-analysis

股票个股分析,实时获取价格涨跌幅,计算技术指标和支撑位,识别缺口并判断支撑压力,智能预测未来3天走势并给出操作建议

Unverified3,646

generate-rag-dataset

Generate a synthetic evaluation dataset from your RAG knowledge base. Creates diverse Q&A pairs with expected answers and relevant context, ready for LangWatch experiments and…

Unverified3,365

data-analytics

Create data pipeline and analytics architecture diagrams using PlantUML syntax with database/analytics stencil icons. Best for ETL pipelines, data lakes, real-time streaming…

Unverified3,067

uv-package-manager

Master the uv package manager for fast Python dependency management, virtual environments, and modern Python project workflows. Use when setting up Python projects, managing…

Unverified2,854

literature-review

帮助用户撰写高质量的文献综述类论文。提供从选题、文献检索、评估筛选、结构规划到最终写作的全流程指导。适用于需要撰写独立文献综述论文或学术论文中文献综述部分的用户。

Unverified2,854

StatsPAI_skill

Use when the user asks to run a full empirical / causal analysis in Python — by default in the style of an applied economics paper (AER / QJE / JPE / ReStud / AEJ) with DID / RD…

Unverified2,854

python-econ-computing

Use when writing Python code for DSGE models, HANK models, numerical economic computation, causal inference, or quantitative economic data analysis

Unverified2,854

ai-ml-skills

27 ai & machine learning skills. Trigger: ML experiments, model training, deep learning, NLP, computer vision. Design: covers frameworks, benchmarks, paper reproduction, and AI…

Unverified2,854

ai-model-benchmarking

Benchmark AI models across 60+ academic evaluation suites and metrics

Unverified2,854

ai-scientist-v2-guide

Automated scientific discovery via agentic tree search by Sakana AI

Unverified2,854

anomaly-detection-papers-guide

Industrial anomaly detection methods and benchmark papers

Unverified2,854

anystyle-api

Citation reference parser using machine learning

Unverified2,854

architecture-design

Use only when creating new registrable ML components that require Factory or Registry patterns.

Unverified2,854

auto-review-loop-minimax

Autonomous multi-round research review loop using MiniMax API. Use when you want to use MiniMax instead of Codex MCP for external review. Trigger with \"auto review loop…

Unverified2,854

bibliography-management-guide

Manage references with BibLaTeX, natbib, and LaTeX bibliography styles

Unverified2,854

bibtex-management-guide

Clean, format, deduplicate, and manage BibTeX bibliography files for LaTeX

Unverified2,854

bokeh-visualization-guide

Guide to Bokeh for interactive browser-based research visualizations

Unverified2,854

causal-inference-guide

Causal inference methods including DiD, IV, RDD, and synthetic control

Unverified2,854

causal-ml

Reference for semiparametric ML estimators: DML with cross-fitting, generalized random forests, debiased regularization, and nuisance function approximation. Covers…

Unverified2,854

check-env

Verifies required tools (Quarto, uv, Python, R, Stata, TeX) and Jupyter kernels are installed. Use when setting up or troubleshooting.

Unverified2,854

citation-network-builder

Build and analyze citation networks from academic reference data

Unverified2,854

codebook

Auto-generates a Markdown codebook from a dataset (CSV, DTA, Excel, Parquet) with types and summary statistics. Use when documenting variables.

Unverified2,854

codebook-pass

调查数据清洗Skill。处理调查数据(CGSS/CHIP/CSS等)时的标准化清洗流程,包括缺失值处理、变量编码统一、数据异常值检测。触发词:数据清洗/调查数据/codebook/数据清洗流程/问卷数据处理

Unverified2,854

computational-chemistry-guide

DFT, molecular simulation, and reaction prediction tools for chemists

Unverified2,854

conciseness-editing-guide

Eliminate wordiness and redundancy in academic prose for clarity

Unverified2,854

cross-disciplinary-ideation

Field connection mapping and systematic ideation for method transfer

Unverified2,854

csv-data-analyzer

Load, explore, clean, and analyze CSV data with statistical summaries

Unverified2,854

data-anomaly-detection

Detect anomalies and outliers in research data using statistical methods

Unverified2,854

data-cleaning-pipeline

Systematic data cleaning workflows for research datasets

Unverified2,854

data-cog-guide

Upload messy CSVs with minimal prompting for deep automated analysis

Unverified2,854

data-scientist

Rigorous data science methodology and mindset for Python research. Covers EDA, data validation, transformation verification, documentation standards, visualization design…

Unverified2,854

dataset-finder-guide

Search and download research datasets from Kaggle, HuggingFace, and repos

Unverified2,854

dataviz-skills

14 data visualization skills. Trigger: charts, plots, figures, publication-quality graphics. Design: one skill per tool with code templates and academic formatting conventions.

Unverified2,854

document-skills

10 document processing skills. Trigger: extracting text from PDFs, parsing references, document Q&A. Design: parsing pipelines (GROBID, marker) and structured extraction tools.

Unverified2,854

env-snapshot

Captures tool versions, packages, and kernel info as a reproducibility record in notes/. Use when documenting the environment.

Unverified2,854

experimental-design-guide

Design rigorous experiments using DOE, factorial designs, and response surfaces

Unverified2,854

geopandas

geopandas spatial data library for Python: manipulation, analysis, and visualization of geographic data. Covers GeoDataFrames, spatial joins, CRS/projections, vector operations…

Unverified2,854

geospatial-viz-guide

Create maps, choropleths, and spatial data visualizations for research

Unverified2,854

google-colab-guide

Run and manage Google Colab notebooks for Python and ML research

Unverified2,854

grad-school-guide

Practical advice for thriving in PhD programs and academic research

Unverified2,854

grobid-pdf-parsing

Extract structured text, metadata, and references from academic PDFs

Unverified2,854

huggingface-inference-guide

Run NLP and CV model inference via Hugging Face free-tier API

Unverified2,854

innovation-management-guide

Innovation metrics, R&D management research, and technology forecasting

Unverified2,854

interactive-viz-guide

Interactive data visualization with Plotly, ECharts, and D3

Unverified2,854

iv-regression-guide

Apply instrumental variables, 2SLS, and address endogeneity issues

Unverified2,854

json-data-visualizer

Guide to JSON Crack for visualizing complex JSON data structures

Unverified2,854

jupyter-notebook-guide

Best practices for computational research notebooks with reproducible workflows

Unverified2,854

kaggle-learner

This skill should be used when the user asks to "learn from Kaggle", "study Kaggle solutions", "analyze Kaggle competitions", or mentions Kaggle competition URLs. Provides access…

Unverified2,854

kedro-pipeline-guide

Build reproducible data science pipelines with Kedro for research projects

Unverified2,854

lens-scholarly-api

Search 300M+ scholarly and patent records via the Lens.org API

Unverified2,854

llm-from-scratch-guide

Build a ChatGPT-like LLM from scratch using PyTorch step by step

Unverified2,854

metabase-analytics-guide

Guide to Metabase for open-source research data analytics and dashboards

Unverified2,854

methods-section-guide

Guide to writing clear and reproducible methodology sections

Unverified2,854

missing-data-handling

Diagnose missing data patterns and apply appropriate imputation strategies

Unverified2,854

ml-causal

This skill covers modern ML-based causal inference methods: Causal Forests (GRF) for heterogeneous treatment effects, Double/Debiased Machine Learning (DML) for partially linear…

Unverified2,854

ml-experiment-tracker

Plan reproducible ML experiment runs with parameters and metrics tracking

Unverified2,854

ml-pipeline-guide

Build and deploy reproducible production ML pipelines for research

Unverified2,854

open-semantic-search-guide

Self-hosted semantic search and text mining platform

Unverified2,854

pandas-data-wrangling

Data cleaning, transformation, and exploratory analysis with pandas

Unverified2,854

papers-we-love-guide

Community-curated directory of influential CS research papers

Unverified2,854

patent-analysis-guide

Patent search, classification, landscape analysis, and prior art mining

Unverified2,854

Full-empirical-analysis-skill

Classical end-to-end empirical analysis workflow in the traditional Python econometric stack — pandas + numpy + scipy + statsmodels + linearmodels + pyfixest + rdrobust + econml…

Unverified2,854

plotly-interactive-guide

Guide to Plotly.py for interactive scientific visualizations in Python

Unverified2,854

polars

Polars DataFrame library for high-performance data manipulation in Python. Covers lazy/eager execution, expressions, I/O (CSV, Parquet, JSON, database), aggregations, joins…

Unverified2,854

prompt-engineering-research

Systematic prompt engineering methods for AI-assisted academic research workf...

Unverified2,854

python-causality-guide

Learn causal inference with Python using the Brave and True handbook

Unverified2,854

python-reproducibility-guide

Reproducible Python environments, notebooks, and literate programming

Unverified2,854

reinforcement-learning-guide

Reinforcement learning fundamentals, algorithms, and research

Unverified2,854

running-causalpy-experiments

Fit, summarize, plot, and interpret a chosen CausalPy experiment. Use after the causal method has been selected, including when configuring PyMC/sklearn models and scale-aware…

Unverified2,854

science-communication

Translating technical data science findings for non-technical audiences. Covers audience analysis, narrative frameworks (Pyramid Principle, SCQA, AIDA), plain-language…

Unverified2,854

scikit-learn

General-purpose machine learning with scikit-learn. Covers unsupervised methods (clustering, GMM, PCA, t-SNE, UMAP, manifold learning, evaluation metrics), supervised methods…

Unverified2,854

species-distribution-guide

Species distribution modeling with MaxEnt, SDM methods, and GBIF data

Unverified2,854

stat-writing

stat-writing

Unverified2,854

statistical-modeling

Use when estimating a statistical or econometric model, running a regression, specifying an identification strategy, testing a hypothesis, or fitting any model to empirical data.…

Unverified2,854

streamline-analyst-guide

End-to-end data analysis AI agent with Streamlit UI

Unverified2,854

survey-data-processing

Clean, recode, and prepare survey response data for analysis

Unverified2,854

svy

svy: design-based analysis of complex survey data in Python. Covers survey design specification (strata, PSU, weights, FPC), variance estimation (Taylor linearization, BRR…

Unverified2,854

tensorflow-guide

TensorFlow best practices for tf.function, GPU memory, and deployment

Unverified2,854

topology-data-analysis

Topological data analysis: persistent homology, Mapper, and TDA tools

Unverified2,854

wrangling-skills

10 data wrangling skills. Trigger: messy data, format conversion, missing values, data reshaping. Design: pipeline-oriented recipes for common data cleaning and transformation…

Unverified2,854

zotero-actions-tags-guide

Zotero workflow automation with custom actions and tags

Unverified2,854

zotero-mcp-guide

Guide to Zotero MCP for connecting Zotero library with AI assistants

Unverified2,854

pyzotero

Interact with Zotero reference management libraries using the pyzotero Python client. Retrieve, create, update, and delete items, collections, tags, and attachments via the…

Unverified2,841

polars

Fast in-memory DataFrame library for datasets that fit in RAM. Use when pandas is too slow but data still fits in memory. Lazy evaluation, parallel execution, Apache Arrow…

Unverified2,841

scikit-learn

Machine learning in Python with scikit-learn. Use when working with supervised learning (classification, regression), unsupervised learning (clustering, dimensionality…

Unverified2,841

aeon

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity…

Unverified2,841

anndata

This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata…

Unverified2,841

arboreto

Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell…

Unverified2,841

bio-admet-prediction

Predicts ADMET properties using ADMETlab 3.0 API or DeepChem models. Estimates bioavailability, CYP inhibition, hERG liability, and 119 toxicity endpoints with uncertainty…

Unverified2,841

bio-atac-seq-atac-qc

Quality control metrics for ATAC-seq data including fragment size distribution, TSS enrichment, FRiP, and library complexity. Use when assessing ATAC-seq library quality before…

Unverified2,841

bio-atac-seq-differential-accessibility

Find differentially accessible chromatin regions between conditions using DiffBind or DESeq2. Use when comparing chromatin accessibility between treatment groups, cell types, or…

Unverified2,841

bio-atac-seq-footprinting

Detect transcription factor binding sites through footprinting analysis in ATAC-seq data using TOBIAS. Use when identifying TF occupancy patterns within accessible regions, as TF…

Unverified2,841

bio-chipseq-motif-analysis

De novo motif discovery and known motif enrichment analysis using HOMER and MEME-ChIP. Identify transcription factor binding motifs in ChIP-seq, ATAC-seq, or other genomic peak…

Unverified2,841

bio-copy-number-cnv-visualization

Visualize copy number profiles, segments, and compare across samples. Create publication-quality plots of CNV data from CNVkit, GATK, or other callers. Use when creating…

Unverified2,841

bio-crispr-screens-base-editing-analysis

Analyzes base editing and prime editing outcomes including editing efficiency, bystander edits, and indel frequencies. Use when quantifying CRISPR base editor results, comparing…

Unverified2,841

bio-crispr-screens-batch-correction

Batch effect correction for CRISPR screens. Covers normalization across batches, technical replicate handling, and batch-aware analysis. Use when combining screens from multiple…

Unverified2,841

bio-crispr-screens-crispresso-editing

CRISPResso2 for analyzing CRISPR gene editing outcomes. Quantifies indels, HDR efficiency, and generates comprehensive editing reports. Use when analyzing amplicon sequencing…

Unverified2,841

bio-crispr-screens-hit-calling

Statistical methods for calling hits in CRISPR screens. Covers MAGeCK, BAGEL2, drugZ, and custom approaches for identifying essential and resistance genes. Use when identifying…

Unverified2,841

bio-crispr-screens-library-design

CRISPR library design for genetic screens. Covers sgRNA selection, library composition, control design, and oligo ordering. Use when designing custom sgRNA libraries for…

Unverified2,841

bio-crispr-screens-mageck-analysis

MAGeCK (Model-based Analysis of Genome-wide CRISPR-Cas9 Knockout) for pooled CRISPR screen analysis. Covers count normalization, gene ranking, and pathway analysis. Use when…

Unverified2,841

bio-crispr-screens-screen-qc

Quality control for pooled CRISPR screens. Covers library representation, read distribution, replicate correlation, and essential gene recovery. Use when assessing screen quality…

Unverified2,841

bio-ctdna-mutation-detection

Detects somatic mutations in circulating tumor DNA using variant callers optimized for low allele fractions with UMI-based error suppression. Reliably detects mutations at VAF…

Unverified2,841

bio-differential-splicing

Detects differential alternative splicing between conditions using rMATS-turbo (BAM-based) or SUPPA2 diffSplice (TPM-based). Reports events with FDR-corrected significance and…

Unverified2,841

bio-epidemiological-genomics-amr-surveillance

Detect and track antimicrobial resistance genes using AMRFinderPlus and ResFinder with epidemiological context. Monitor resistance trends and identify emerging resistance…

Unverified2,841

bio-epidemiological-genomics-pathogen-typing

Perform multi-locus sequence typing (MLST), core genome MLST, and SNP-based strain typing for bacterial isolate characterization using mlst and chewBBACA. Use when identifying…

Unverified2,841

bio-epidemiological-genomics-transmission-inference

Infer pathogen transmission networks and identify likely transmission pairs using TransPhylo and outbreak reconstruction algorithms. Estimate who-infected-whom from genomic and…

Unverified2,841

bio-epidemiological-genomics-variant-surveillance

Assign pathogen lineages and track variants using Nextclade and pangolin for viral surveillance. Monitor variant prevalence and identify emerging variants of concern. Use when…

Unverified2,841

bio-fragment-analysis

Analyzes cfDNA fragment size distributions and fragmentomics features using FinaleToolkit or Griffin. Extracts nucleosome positioning patterns, fragment ratios, and DELFI-style…

Unverified2,841

bio-genome-engineering-off-target-prediction

Predict CRISPR off-target sites using Cas-OFFinder and CFD scoring algorithms. Identify potential unintended cleavage sites genome-wide and assess guide specificity. Use when…

Unverified2,841

bio-hi-c-analysis-compartment-analysis

Detect A/B compartments from Hi-C data using cooltools and eigenvector decomposition. Identify active (A) and inactive (B) chromatin compartments from contact matrices. Use when…

Unverified2,841

bio-hi-c-analysis-contact-pairs

Process Hi-C read pairs using pairtools. Parse alignments, filter duplicates, classify pairs, and generate contact statistics from Hi-C sequencing data. Use when processing raw…

Unverified2,841

bio-hi-c-analysis-hic-data-io

Load, convert, and manipulate Hi-C contact matrices using cooler format. Read .cool/.mcool files, convert from .hic format, access matrix data, and export to different formats.…

Unverified2,841

bio-hi-c-analysis-hic-differential

Compare Hi-C contact matrices between conditions to identify differential chromatin interactions. Compute log2 fold changes, statistical significance, and visualize differential…

Unverified2,841

bio-hi-c-analysis-hic-visualization

Visualize Hi-C contact matrices, TADs, loops, and genomic features using matplotlib, cooltools, and HiCExplorer. Create triangle plots, virtual 4C, and multi-track figures. Use…

Unverified2,841

bio-hi-c-analysis-loop-calling

Detect chromatin loops and point interactions from Hi-C data using cooltools, chromosight, and HiCCUPS-like methods. Identify CTCF-mediated loops and enhancer-promoter contacts.…

Unverified2,841

bio-hi-c-analysis-matrix-operations

Balance, normalize, and transform Hi-C contact matrices using cooler and cooltools. Apply iterative correction (ICE), compute expected values, and generate observed/expected…

Unverified2,841

bio-hi-c-analysis-tad-detection

Call topologically associating domains (TADs) from Hi-C data using insulation score, HiCExplorer, and other methods. Identify domain boundaries and hierarchical domain structure.…

Unverified2,841

bio-imaging-mass-cytometry-cell-segmentation

Cell segmentation from multiplexed tissue images. Covers deep learning (Cellpose, Mesmer) and classical approaches for nuclear and whole-cell segmentation. Use when extracting…

Unverified2,841

bio-imaging-mass-cytometry-data-preprocessing

Load and preprocess imaging mass cytometry (IMC) and MIBI data. Covers MCD/TIFF handling, hot pixel removal, and image normalization. Use when starting IMC analysis from raw MCD…

Unverified2,841

bio-imaging-mass-cytometry-interactive-annotation

Interactive cell type annotation for IMC data. Covers napari-based annotation, marker-guided labeling, training data generation, and annotation validation. Use when manually…

Unverified2,841

bio-imaging-mass-cytometry-phenotyping

Cell type assignment from marker expression in IMC data. Covers manual gating, clustering, and automated classification approaches. Use when assigning cell types to segmented IMC…

Unverified2,841

bio-imaging-mass-cytometry-quality-metrics

Quality metrics for IMC data including signal-to-noise, channel correlation, tissue integrity, and acquisition QC. Use when assessing data quality before analysis or…

Unverified2,841

bio-imaging-mass-cytometry-spatial-analysis

Spatial analysis of cell neighborhoods and interactions in IMC data. Covers neighbor graphs, spatial statistics, and interaction testing. Use when analyzing spatial relationships…

Unverified2,841

bio-immunoinformatics-epitope-prediction

Predict B-cell and T-cell epitopes using BepiPred, IEDB tools, and structure-based methods for vaccine and antibody design. Identify immunogenic regions in antigens. Use when…

Unverified2,841

bio-immunoinformatics-immunogenicity-scoring

Score and prioritize neoantigens and epitopes for immunogenicity using multi-factor models combining MHC binding, processing, expression, and sequence features. Rank candidates…

Unverified2,841

bio-immunoinformatics-mhc-binding-prediction

Predict peptide-MHC class I and II binding affinity using MHCflurry and NetMHCpan neural network models. Identify potential T-cell epitopes from protein sequences. Use when…

Unverified2,841

bio-immunoinformatics-neoantigen-prediction

Identify tumor neoantigens from somatic mutations using pVACtools for personalized cancer immunotherapy. Predict mutant peptides that bind patient HLA and may elicit T-cell…

Unverified2,841

bio-immunoinformatics-tcr-epitope-binding

Predict TCR-epitope specificity using ERGO-II and deep learning models for T-cell receptor antigen recognition. Match TCRs to their cognate epitopes or predict TCR targets. Use…

Unverified2,841

bio-long-read-sequencing-isoseq-analysis

Analyze PacBio Iso-Seq data for full-length isoform discovery and quantification. Use when characterizing transcript diversity or identifying novel splice variants.

Unverified2,841

bio-metabolomics-lipidomics

Specialized lipidomics analysis for lipid identification, quantification, and pathway interpretation. Covers LC-MS lipidomics with LipidSearch, MS-DIAL, and LipidMaps annotation.…

Unverified2,841

bio-metabolomics-metabolite-annotation

Metabolite identification from m/z and retention time. Covers database matching, MS/MS spectral matching, and confidence level assignment. Use when assigning compound identities…

Unverified2,841

bio-metabolomics-msdial-preprocessing

MS-DIAL-based metabolomics preprocessing as alternative to XCMS. Covers peak detection, alignment, annotation, and export for downstream analysis. Use when processing MS-DIAL…

Unverified2,841

bio-metabolomics-targeted-analysis

Targeted metabolomics analysis using MRM/SRM with standard curves. Covers absolute quantification, method validation, and quality assessment. Use when quantifying specific…

Unverified2,841

bio-metagenomics-abundance

Species abundance estimation using Bracken with Kraken2 output. Redistributes reads from higher taxonomic levels to species for more accurate estimates. Use when accurate…

Unverified2,841

bio-metagenomics-functional-profiling

Profile functional potential of metagenomes using HUMAnN3 and similar tools. Use when obtaining pathway abundances, gene family counts, or functional annotations from metagenomic…

Unverified2,841

bio-metagenomics-kraken

Taxonomic classification of metagenomic reads using Kraken2. Fast k-mer based classification against RefSeq database. Use when performing initial taxonomic classification of…

Unverified2,841

bio-metagenomics-metaphlan

Marker gene-based taxonomic profiling using MetaPhlAn 4. Provides accurate species-level relative abundances using clade-specific markers. Use when accurate taxonomic profiling…

Unverified2,841

bio-metagenomics-strain-tracking

Track bacterial strains using MASH, sourmash, fastANI, and inStrain. Compare genomes, detect contamination, and monitor strain-level variation. Use when needing sub-species…

Unverified2,841

bio-metagenomics-visualization

Visualize metagenomic profiles using R (phyloseq, microbiome) and Python (matplotlib, seaborn). Create stacked bar plots, heatmaps, PCA plots, and diversity analyses. Use when…

Unverified2,841

bio-methylation-based-detection

Analyzes cfDNA methylation patterns for cancer detection using cfMeDIP-seq or bisulfite sequencing with MethylDackel. Identifies cancer-specific methylation signatures and…

Unverified2,841

bio-methylation-calling

Extract methylation calls from Bismark BAM files using bismark_methylation_extractor. Generates per-cytosine reports for CpG, CHG, and CHH contexts. Use when extracting…

Unverified2,841

bio-microbiome-functional-prediction

Predict metagenome functional content from 16S rRNA marker gene data using PICRUSt2. Infer KEGG, MetaCyc, and EC abundances from ASV tables. Use when functional profiling is…

Unverified2,841

bio-molecular-descriptors

Calculates molecular descriptors and fingerprints using RDKit. Computes Morgan fingerprints (ECFP), MACCS keys, Lipinski properties, QED drug-likeness, TPSA, and 3D conformer…

Unverified2,841

bio-orchestrator

Meta-agent that routes bioinformatics requests to specialised sub-skills. Handles file type detection, analysis planning, report generation, and reproducibility export.

Unverified2,841

bio-proteomics-data-import

Load and parse mass spectrometry data formats including mzML, mzXML, and quantification tool outputs like MaxQuant proteinGroups.txt. Use when starting a proteomics analysis with…

Unverified2,841

bio-proteomics-dia-analysis

Data-independent acquisition (DIA) proteomics analysis with DIA-NN and other tools. Use when analyzing DIA mass spectrometry data with library-free or library-based workflows for…

Unverified2,841

bio-proteomics-differential-abundance

Statistical testing for differentially abundant proteins between conditions. Covers limma and MSstats workflows with multiple testing correction. Use when identifying proteins…

Unverified2,841