AIAny
Icon for item

Hermes Function-Calling V1

Structured dataset for training and evaluating LLM agentic behavior: function-calling conversations, JSON-mode structured outputs, and extraction samples for teaching models to generate tool calls and strict structured responses. Includes single-turn and multi-turn scenarios across several configs.

Introduction

Oops! Something went wrong

[next-mdx-remote-client] error compiling MDX: Expected a closing tag for `<tool_response>` (6:74-6:89) before the end of `paragraph` 4 | ## What sets it apart 5 | - Multi-config mix: five dataset configs (func_calling_singleturn, func_calling, glaive_func_calling, json_mode_agentic, json_mode_singleturn) that cover single-turn calls, multi-turn agentic dialogues, cleaned glaive samples, and advanced JSON-mode responses. This variety helps models learn both simple and complex function-calling patterns. > 6 | - Explicit tool-call format: examples use XML-like `<tools>`/<tool_call>/<tool_response> conventions and pydantic-style schemas so models learn to emit machine-parseable function calls and to parse tool responses back into natural language or structured outputs. | ^ 7 | - Focus on structured extraction and agent skills: includes JSON-mode agentic samples and schema-driven extraction tasks useful for building reliable agents, RAG pipelines, and function-calling integrations. 8 | - Dataset scale & license: ~11,578 examples (size category 10K–100K), Apache-2.0 license, prepared for Hugging Face Datasets consumption with common libraries like pandas and polars. More information: https://mdxjs.com/docs/troubleshooting-mdx

Information

  • Websitehuggingface.co
  • OrganizationsNousResearch
  • Authorsinterstellarninja, Teknium, THEODOROS
  • Published date2024/08/14

Categories

More Items

Hugging Face

Provides 315,000 pairwise human-preference votes comparing 15 English TTS models over 300 operational prompts, with 4,500 high‑quality audio renders and structured vote/pair/prompt records for training or evaluating preference/reward models. Metadata under CC-BY-4.0; audio use governed by model providers' terms.

Hugging Face

A 10‑billion‑document retrieval benchmark with per‑document 768‑dim unit‑norm dense embeddings and mGTE sparse embeddings, FineWeb text/metadata, and exact top‑1000 MS MARCO ground truth for ~120k queries. Built for large‑scale evaluation of dense/sparse/hybrid retrieval, filtered search, indexing, ANNS algorithms, and embedding compression.

Hugging Face

Provides 4.5 billion TikTok video records with captions, timestamps, music IDs and engagement counts for research; split across 27 zstd-compressed Parquet files (~289 GB) and sampled via TikTok's mobile API; released for research-use only with privacy and ToS caveats.