✦
Resource profile / Dataset Viewer Workflows
About this skill

Workflow & requirements

Hugging Face Dataset Viewer

Use this skill to execute read-only Dataset Viewer API calls for dataset exploration and extraction.

Core workflow

  1. Optionally validate dataset availability with /is-valid.
  2. Resolve config + split with /splits.
  3. Preview with /first-rows.
  4. Paginate content with /rows using offset and length (max 100).
  5. Use /search for text matching and /filter for row predicates.
  6. Retrieve parquet links via /parquet and totals/metadata via /size and /statistics.

Defaults

  • Base URL: https://datasets-server.huggingface.co
  • Default API method: GET
  • Query params should be URL-encoded.
  • offset is 0-based.
  • length max is usually 100 for row-like endpoints.
  • Gated/private datasets require Authorization: Bearer <HF_TOKEN>.

Dataset Viewer

  • Validate dataset: /is-valid?dataset=<namespace/repo>
  • List subsets and splits: /splits?dataset=<namespace/repo>
  • Preview first rows: /first-rows?dataset=<namespace/repo>&config=<config>&split=<split>
  • Paginate rows: /rows?dataset=<namespace/repo>&config=<config>&split=<split>&offset=<int>&length=<int>
  • Search text: /search?dataset=<namespace/repo>&config=<config>&split=<split>&query=<text>&offset=<int>&length=<int>
  • Filter with predicates: /filter?dataset=<namespace/repo>&config=<config>&split=<split>&where=<predicate>&orderby=<sort>&offset=<int>&length=<int>
  • List parquet shards: /parquet?dataset=<namespace/repo>
  • Get size totals: /size?dataset=<namespace/repo>
  • Get column statistics: /statistics?dataset=<namespace/repo>&config=<config>&split=<split>
  • Get Croissant metadata (if available): /croissant?dataset=<namespace/repo>

Pagination pattern:

curl "https://datasets-server.huggingface.co/rows?dataset=stanfordnlp/imdb&config=plain_text&split=train&offset=0&length=100"
curl "https://datasets-server.huggingface.co/rows?dataset=stanfordnlp/imdb&config=plain_text&split=train&offset=100&length=100"

When pagination is partial, use response fields such as num_rows_total, num_rows_per_page, and partial to drive continuation logic.

Search/filter notes:

  • /search matches string columns (full-text style behavior is internal to the API).
  • /filter requires predicate syntax in where and optional sort in orderby.
  • Keep filtering and searches read-only and side-effect free.

For CLI-based parquet URL discovery or SQL, use the hf-cli skill with hf datasets parquet and hf datasets sql.

Creating and Uploading Datasets

Use one of these flows depending on dependency constraints.

Zero local dependencies (Hub UI):

  • Create dataset repo in browser: https://huggingface.co/new-dataset
  • Upload parquet files in the repo "Files and versions" page.
  • Verify shards appear in Dataset Viewer:
curl -s "https://datasets-server.huggingface.co/parquet?dataset=<namespace>/<repo>"

Low dependency CLI flow (npx @huggingface/hub / hfjs):

  • Set auth token:
export HF_TOKEN=<your_hf_token>
  • Upload parquet folder to a dataset repo (auto-creates repo if missing):
npx -y @huggingface/hub upload datasets/<namespace>/<repo> ./local/parquet-folder data
  • Upload as private repo on creation:
npx -y @huggingface/hub upload datasets/<namespace>/<repo> ./local/parquet-folder data --private

After upload, call /parquet to discover <config>/<split>/<shard> values for querying with @~parquet.

Agent Traces

The Hub supports raw agent session traces from Claude Code, Codex, and Pi Agent. Upload them to Hugging Face Datasets as original JSONL files and the Hub can auto-detect the trace format, tag the dataset as Traces, and enable the trace viewer for browsing sessions, turns, tool calls, and model responses. Common local session directories:

  • Claude Code: ~/.claude/projects
  • Codex: ~/.codex/sessions
  • Pi: ~/.pi/agent/sessions

Default to private dataset repos because traces can contain prompts, file paths, tool outputs, secrets, or PII. Preserve the raw .jsonl files and nest them by project/cwd instead of uploading every session at the dataset root.

hf repos create <namespace>/<repo> --type dataset --private --exist-ok
hf upload <namespace>/<repo> ~/.codex/sessions codex/<project-or-cwd> --type dataset
PACKAGE TRANSPARENCY

Inspect before installing

Source: Hugging Face · Apache-2.0 · SHA-256 shown alongside the download.

5 files11256 ZIP bytes0 script/code files

License file included. A license and checksum are not a security certification. Review package instructions and scripts before running them.

View files and uncompressed sizes
Machine-readable installation guide →
CATALOG REVIEW NOTES

Know what you need before installing

Source and packaging checks recorded on 2026-10-03. These notes are not safety certification or measured task performance.

Requirements

An HTTP-capable client; internet access; HF_TOKEN for gated or private datasets.

Costs, access & practical limits

Downloaded datasets have their own licenses and may contain personal data. Public metadata access does not grant unrestricted reuse.

View the recorded checks
  • Pinned upstream source
  • Original license and notices preserved
  • Archive paths and metadata validated
  • Local Markdown reference links checked
  • Python files syntax-checked where present

Upstream commit: ca0325bb20b2d0a1b2efa893670c4c72f79e707b

Runtime status: not tested by this catalog. Configure your client and test the skill in your own environment.

SCENARIOS

Inputs, criteria and recorded outcomes

Records are supplied by the site administrator and bound to a specific package. They are not third-party safety certification. This page does not execute skills.

Dataset Viewer endpoint compatibility (read-only)

Reported passed · vca0325bb20b2

View input and acceptance criteria

Input

Use public dataset lhoestq/demo1. Request split metadata, one row at offset 0, and size metadata using the documented Dataset Viewer HTTP endpoints. Do not save row values or use credentials.

Acceptance criteria

All three requests return HTTP 200; splits are non-empty; the row response contains one entry; the size response includes a size object. Scope: endpoint compatibility, not an agent end-to-end evaluation.

Recorded outcome

{
  "scope": "Read-only HTTP endpoint compatibility smoke test; not an end-to-end agent evaluation.",
  "dataset": "lhoestq/demo1",
  "requests": [
    {
      "url": "https://datasets-server.huggingface.co/splits?dataset=lhoestq%2Fdemo1",
      "status": 200,
      "elapsed_seconds": 0.797
    },
    {
      "url": "https://datasets-server.huggingface.co/rows?dataset=lhoestq%2Fdemo1&config=default&split=train&offset=0&length=1",
      "status": 200,
      "elapsed_seconds": 0.953
    },
    {
      "url": "https://datasets-server.huggingface.co/size?dataset=lhoestq%2Fdemo1",
      "status": 200,
      "elapsed_seconds": 1.75
    }
  ],
  "row_count_checked": 1,
  "row_values_saved": false,
  "environment": "3.12.4 / Python urllib, no token, no model invocation",
  "result": "passed",
  "executed_at": "2026-10-02T16:42:24.023840+00:00"
}

Environment

3.12.4 / Python urllib, no token, no model invocation; Windows local workstation; no upstream script execution. Scope: HTTP endpoint smoke test only.

Package SHA-256: 26311310bac44188a2fae8afa8ea92fbe8eac5092aa31ed1bc3e99d8b27d4566

Outcome recorded: 2026-10-02 16:42 UTC

Community reviews

★ New

Be the first to share your experience.

Sign in to leave a review →

More to explore

View all ↗