AIAny

Tag

Explore by tags

Hugging Face

Contains tech-blog posts scraped from Habr (primarily Russian, some English) in Parquet format with ~100K–1M records. Suited for multilingual text-generation and language-model fine-tuning; license is not specified, so verify before redistribution.

Hugging Face

Provides 52,000 English instruction–response pairs generated by OpenAI's text-davinci-003 for instruction-tuning language models. Released under CC BY-NC 4.0; low-cost synthetic data useful for research but contains model-generated biases and errors.

Hugging Face

A 15,000+ English instruction–response corpus for fine-tuning and evaluating LLM instruction-following behavior. Contains human-authored prompts and answers across categories (closed/open QA, summarization, extraction, classification, brainstorming) and is released under CC BY-SA 3.0.

Hugging Face

Provides manually curated Japanese instruction pairs (questions and safe reference answers) for improving LLM output safety, covering broad harm categories and regionally sensitive cases. Includes English meta-tags and standard splits for benchmarking and fine-tuning.

Hugging Face

Structured dataset for training and evaluating LLM agentic behavior: function-calling conversations, JSON-mode structured outputs, and extraction samples for teaching models to generate tool calls and strict structured responses. Includes single-turn and multi-turn scenarios across several configs.

Hugging Face

A mixed instruction dataset for SFT and RLHF research that combines chat, math, code and instruction-following samples from multiple public datasets under an Apache-2.0-compatible license; intended for instruction tuning and evaluation.

Hugging Face

A curated dataset of ~30,000 CUDA kernels generated by an agentic pipeline, including reference PyTorch implementations, runtime metrics, NCU/Torch/Clang-Tidy profiles, error messages and correctness labels — released under CC-BY-4.0 for model fine-tuning and offline RL/optimization research.

Hugging Face

Provides 1.7M+ synthetic and real infographic charts paired with their tabular data for training and evaluating multimodal models on infographic understanding, chart-to-table extraction, chart code generation, and example-based chart synthesis.

Hugging Face

Provides MS MARCO queries, passages and answers translated into 14 Indic languages while keeping the original English content and per-example translation metadata. Includes train/validation splits, passage selection flags, and translation model parameters for multilingual IR, QA and RAG research.

Hugging Face

Collects ~200,000 human responses to 20 visual/semantic association questions (e.g., Bouba–Kiki), with per-response image options and demographic metadata — useful for cross‑cultural perception and evaluation of multimodal systems, but not guaranteed as a rigorously controlled experimental sample.

Hugging Face

Provides 100 real-world, open-ended research tasks paired with expert-written rubrics (around 40 weighted criteria per task) to evaluate long-form, web-browsing research agents on factual accuracy, analysis depth, presentation, and citation quality.

Hugging Face

Provides 6,000 runnable, operator-level PyTorch tasks for training and evaluating CUDA kernel generation models; each sample includes executable code, operator descriptors, and provenance tags, with execution-driven filtering to ensure reproducibility and contamination control.