BB SKILLS / LEARN

Analyze local CSV data with DuckDB agent skills

DuckDB skills can help an agent turn a data question into SQL against files on your own machine. The result still needs checking: start with a small, non-sensitive fixture, inspect the generated query, and compare its output with a calculation you can verify yourself.

This walkthrough uses our independently packaged DuckDB Local Data Queries and DuckDB Database Sessions resources. Both preserve the original DuckDB Foundation MIT notice and the exact upstream instructions. BB Skills adds execution notes and, for database sessions, a Python helper that quotes paths and aliases when appending state. These ZIPs are individual skills, not the complete upstream plugin.

What you need

Use a macOS or Linux environment with DuckDB CLI and Bash. The database-session helper also needs Python 3 with POSIX file locking. Our recorded CLI checks used DuckDB 1.5.6 on Linux; they do not establish Windows agent compatibility. Commands that reference other /duckdb-skills:* skills require the complete upstream Claude Code plugin. Install dependencies deliberately and inspect downloaded files before running them.

The small local example below needs no remote storage, model API or extension download. An AI client may still charge for its model usage. BB Skills does not run third-party skills online.

Make a five-row fixture

Save this as sales.csv in a scratch directory:

category,amount,note
books,10,first
books,20,second
games,5,third
games,15,fourth
music,25,fifth

There are five rows. The category totals are books 30, games 20 and music 25, giving a total of 75. This known answer makes the exercise useful for detecting a query that runs successfully but answers the wrong question.

Bound file access and output

Run this from the directory containing sales.csv. It permits that file, disables other external access and persistent secrets, then locks the configuration for the query session:

duckdb :memory: -csv <<'SQL'
SET allowed_paths=['sales.csv'];
SET enable_external_access=false;
SET allow_persistent_secrets=false;
SET lock_configuration=true;
SELECT category, sum(amount) AS total
FROM 'sales.csv'
GROUP BY ALL
ORDER BY category;
SQL

The expected output is:

category,total
books,30
games,20
music,25

Inspect SQL before execution. A quoted heredoc preserves shell-like text as SQL input rather than expanding it in the shell. For an absolute path containing a single quote, double that quote inside the SQL string literal. Identifiers such as table names need double quotes, with embedded double quotes doubled. Avoid copying an untrusted path into a shell command or interpolating raw SQL into a double-quoted -c string.

These settings limit DuckDB file access; they are not a complete process sandbox. For unfamiliar data, use a disposable environment with separate memory, CPU and time limits. Keep queries bounded with aggregation or LIMIT before returning results to an agent conversation.

Turn it into a useful agent request

A suitable first request is: "Read only this synthetic CSV. Count the rows and calculate the amount total by category. Show the SQL before running it. Return the three category totals and compare them with books 30, games 20 and music 25. Do not install extensions or access other files."

This is a test prompt, not a promise that every client will follow it. Review the requested tool actions and verify the output. With real data, also check schema inference, NULL values, duplicate rows, encoding, amount units and whether the grouping matches the question.

Restore a database session carefully

The database-session resource inspects a trusted DuckDB file and appends an attachment to state.sql. Read existing state before using duckdb -init: state can contain arbitrary SQL, extension loads or secrets. Keep it private and outside version control. Session mode restores a user-trusted database and has different access assumptions from the restricted CSV example above.

The packaged helper preserves prior state, quotes the database path and alias, rejects a conflicting alias and symbolic-link state file, and sets the state file to mode 0600. It writes attachment SQL; it does not validate or execute the rest of the state file. Check the resource's scoped scenario record for the exact version and cases we observed.

What our records establish

We ran 28 selected assertions in a network-disabled, disposable Linux container with synthetic data, 512 MiB memory and one CPU. Thirteen cover selected query behavior and access restrictions. Fifteen cover database validation, schema inspection, session restoration and the packaged state helper. Each public record is bound to the tested package checksum and source file hashes.

These checks do not measure an AI model's ability to choose correct SQL, production performance, large-data behavior, all upstream features, malicious database resistance or every permission boundary. The packages retain a source-review status; a bounded CLI observation does not certify the entire skill.

Primary references

For corrections to this walkthrough or a resource, contact [email protected] with its page URL.