tao-mine-nearest-neighbors

v2026.09.24

Run TAO Data Services TMM nearest-neighbor mining from embedding parquet files. Use when a workflow needs to mine source samples closest to target samples.

GitHub
Install command
npx skhub add nvidia/tao-mine-nearest-neighbors
Markdown
SKILL.md

TAO Mine Nearest Neighbors

Use this skill to run TAO Data Services TMM nearest-neighbor mining. The skill consumes embedding parquets and writes a mined source-sample parquet plus a mining summary. It does not compute embeddings; upstream steps must produce the source and target embedding parquets first.

The container entrypoint is:

tmm nearest_neighbors -e /absolute/path/to/nearest_neighbors.yaml

TAO Data Services requires -e/--experiment_spec_file. The tmm console script converts that YAML into Hydra --config-path and --config-name arguments internally.

Inputs

The user can provide either an existing nearest-neighbors YAML spec or the fields needed to generate one.

Required spec fields:

FieldMeaning
source_parquetAbsolute path to the candidate/source embeddings parquet.
target_parquetAbsolute path to the target/query embeddings parquet.
output_parquetAbsolute path where TAO Data Services should write mined source filepaths.

Common optional fields:

FieldDefaultMeaning
topn5Number of nearest source samples to retrieve per target sample.
knn_metriccosineOne of cosine, euclidean, or manhattan.
source_embed_column_nameembeddingEmbedding column in source_parquet.
target_embed_column_nameembeddingEmbedding column in target_parquet.
filter_by_label"false"String flag. When "true", TAO DS filters neighbors by matching label columns when both parquets provide labels.
distance_threshold-1.0Maximum distance to keep. Negative disables thresholding.

Both input parquets must contain a filepath column and a list-like embedding column. If filter_by_label is "true", both parquets should also contain label.

The default template is assets/default_nearest_neighbors.yaml.

Quick Start

Run from the tao-skill-bank repo root. Resolve the pinned TAO Data Services image from versions.yaml, verify the spec, mount the run root with identical host/container paths, and stream the Docker logs.

SPEC=/absolute/path/to/nearest_neighbors.yaml
RUN_ROOT=/absolute/path/that/contains/specs/data/and/results
GPU_COUNT=1

python3 skills/data/tao-mine-nearest-neighbors/scripts/verify_nearest_neighbors_spec.py \
  --spec "$SPEC"

DS_IMAGE="$(scripts/resolve_versions_key.py images.tao_toolkit.data_services)"

docker run --rm --gpus "$GPU_COUNT" --shm-size=8g --network=host \
  -v "$RUN_ROOT:$RUN_ROOT" \
  -w "$RUN_ROOT" \
  "$DS_IMAGE" \
  tmm nearest_neighbors -e "$SPEC"

Use at least one GPU. Choose GPU_COUNT from the hardware available to the host or platform that will run the container. If the user does not know the right value, inspect the host with nvidia-smi -L or ask which GPU allocation the run should use.

Do not pass --user $(id -u):$(id -g) to the TAO data-services container unless you have verified the image supports that UID. Some TAO DS images import Python packages that call getpass.getuser() at startup and fail when the UID is not present in /etc/passwd.

Generate A Spec

If the user provides source/target/output parquet paths instead of a ready spec, generate a spec from the default template:

python3 skills/data/tao-mine-nearest-neighbors/scripts/prepare_nearest_neighbors_spec.py \
  --source-parquet /absolute/path/source_embeddings.parquet \
  --target-parquet /absolute/path/target_embeddings.parquet \
  --output-parquet /absolute/path/results/mined.parquet \
  --output-spec /absolute/path/specs/nearest_neighbors.yaml \
  --topn 5 \
  --knn-metric cosine \
  --filter-by-label false \
  --distance-threshold -1.0

The generated YAML uses absolute paths. Keep the spec, input parquets, and output directory under RUN_ROOT so the same paths resolve inside the container.

Preflight

Before launching Docker:

  1. Verify Docker and GPU access:
docker info > /dev/null
nvidia-smi -L
  1. Resolve and pull the data-services image if needed:
DS_IMAGE="$(scripts/resolve_versions_key.py images.tao_toolkit.data_services)"
docker image inspect "$DS_IMAGE" > /dev/null || docker pull "$DS_IMAGE"
  1. Validate the spec:
python3 skills/data/tao-mine-nearest-neighbors/scripts/verify_nearest_neighbors_spec.py \
  --spec "$SPEC"
  1. Confirm RUN_ROOT contains the spec, both input parquets, and the output directory. Mount RUN_ROOT to the same absolute path inside Docker.

Outputs

The skill promises the artifacts named by the spec:

ArtifactLocation
mined parquetoutput_parquet
mining summarymining_summary.txt next to output_parquet

The current TAO Data Services nearest_neighbors task writes a mined parquet with unique source filepath rows. The summary file reports mining counts such as queries processed, neighbors considered, duplicates removed, and any label/distance filtering.

Troubleshooting

The subtask nearest_neighbors requires -e/--experiment_spec_file: rerun with tmm nearest_neighbors -e "$SPEC". Hydra overrides alone are not enough.

Input parquet not found inside Docker: the YAML path must be visible inside the container. Use a RUN_ROOT mount where the host and container paths are identical.

Output directory is not writable after Docker exits: the TAO DS container may have written files as root. Inform the user, report which artifacts were produced, and ask whether to repair permissions on the output directory before continuing.

No GPU or cuDF/cuML errors: nearest-neighbor mining requires at least one CUDA GPU. Check nvidia-smi -L, the Docker --gpus flag, and the NVIDIA container toolkit installation.

Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

Apache-2.0

Source path

skills/tao-mine-nearest-neighbors

Default branch

main

Latest commit

ef46204

Tree SHA

94ca43b