AIAny
Icon for item

Nemotron-SFT-SWE-v3.5

Provides agentic instruction‑tuning trajectories for software‑engineering tasks, formatted for supervised fine‑tuning and agent training. Contains multi‑file edits, tests, docs and structured agent traces (≈5,115 records, 1.9 GiB). Intended for commercial use; licensed CC‑BY 4.0 with additional permissive licenses.

Introduction

Most code-focused SFT blends target single-shot repairs or test generation; this dataset emphasizes agentic, multi-step workflows and repository-aware edits so models learn planning, tool use, and cross-file reasoning. It packages structured agent traces collected with the OpenCode harness and curated for supervised fine‑tuning of software-engineering agents.

What Sets It Apart
  • Agentic trajectories rather than isolated Q&A: includes stepwise agent actions, tool calls and patch-style edits so fine‑tuned models can learn multi-step workflows and intermediate state management (so what: better support for agents that must plan and execute sequences across files).
  • Multi-artifact, multi-file focus: examples include source, tests, docs and config changes together, not just single-file patches (so what: improves models’ repository-aware reasoning and regression‑free edits).
  • Compact, curated SFT target: ~5,115 records (1.9 GiB) designed for supervised fine‑tuning rather than large‑scale pretraining (so what: faster iteration for teams training moderate-size models or distillation pipelines).
  • Clear commercial-use signal and mixed permissive licensing: distributed under CC‑BY 4.0 with additional permissive licenses noted (so what: suitable for product integration but check organiational legal requirements).
Who It's For and Tradeoffs

Great fit if you are training or distilling LLMs to act as autonomous coding agents, need cross-file repair/test-generation examples, or want agent traces that reflect tool use and stepwise reasoning. Look elsewhere if you need very large-scale SWE corpora (tens or hundreds of thousands of examples) for pretraining, raw repository snapshots, or if your project policies prohibit the dataset's commercial-use terms. In short: efficient SFT material for repository-aware agent behavior, with explicit tradeoffs around scale and licensing.

More Items

Hugging Face

Provides 315,000 pairwise human-preference votes comparing 15 English TTS models over 300 operational prompts, with 4,500 high‑quality audio renders and structured vote/pair/prompt records for training or evaluating preference/reward models. Metadata under CC-BY-4.0; audio use governed by model providers' terms.

Hugging Face

A 10‑billion‑document retrieval benchmark with per‑document 768‑dim unit‑norm dense embeddings and mGTE sparse embeddings, FineWeb text/metadata, and exact top‑1000 MS MARCO ground truth for ~120k queries. Built for large‑scale evaluation of dense/sparse/hybrid retrieval, filtered search, indexing, ANNS algorithms, and embedding compression.

Hugging Face

Provides 4.5 billion TikTok video records with captions, timestamps, music IDs and engagement counts for research; split across 27 zstd-compressed Parquet files (~289 GB) and sampled via TikTok's mobile API; released for research-use only with privacy and ToS caveats.