Skip to content
6 changes: 5 additions & 1 deletion docs/docs.json
Original file line number Diff line number Diff line change
Expand Up @@ -67,7 +67,7 @@
"group": "Model training",
"pages": [
"training/why-lancedb",
"training/index",
"training/data-loading",
"training/torch",
"training/object-detection",
"training/vlm-finetuning",
Expand Down Expand Up @@ -573,6 +573,10 @@
{
"source": "/huggingface/datasets",
"destination": "/datasets"
},
{
"source": "/training",
"destination": "/training/data-loading"
}
]
}
2 changes: 1 addition & 1 deletion docs/index.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,7 @@ feature engineering, search and retrieval, and efficient training data access.
<Card title="Train and fine-tune models" icon="fire" href="/training/why-lancedb">
Learn why LanceDB works well as the data layer for training workloads.
</Card>
<Card title="Load data into PyTorch" icon="boxes-stacked" href="/training/">
<Card title="Load data into PyTorch" icon="boxes-stacked" href="/training/data-loading">
Use LanceDB tables and permutations for projected, shuffled, random-access training reads.
</Card>
<Card title="Browse ready-to-use datasets" icon="database" href="/datasets">
Expand Down
2 changes: 1 addition & 1 deletion docs/quickstart.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -466,7 +466,7 @@ updates, versioning, indexing, full-text search, hybrid search, and reranking.
<Card
title="Data loading and shuffles"
icon="boxes-stacked"
href="/training/"
href="/training/data-loading"
>
Use LanceDB for projected, shuffled, random-access reads in training workflows.
</Card>
Expand Down
475 changes: 475 additions & 0 deletions docs/training/data-loading.mdx

Large diffs are not rendered by default.

512 changes: 0 additions & 512 deletions docs/training/index.mdx

This file was deleted.

4 changes: 2 additions & 2 deletions docs/training/object-detection.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ icon: car

This example walks through fine-tuning an autonomous vehicle (AV) perception model on targeted failure-mode slices of [BDD100K](https://www.bdd100k.com/) — riders, nighttime pedestrians, and distant pedestrians — using LanceDB as a single multimodal table from raw JPEG bytes through to the PyTorch training loop.

The full pipeline lives in the [lancedb/training](https://github.com/lancedb/training/tree/main/object-detection) repository. This page focuses on the parts most relevant to training: defining curated splits as materialized views, loading them through the [`Permutation`](/training/) API, and pinning checkpoints to an exact data version.
The full pipeline lives in the [lancedb/training](https://github.com/lancedb/training/tree/main/object-detection) repository. This page focuses on the parts most relevant to training: defining curated splits as materialized views, loading them through the [`Permutation`](/training/data-loading) API, and pinning checkpoints to an exact data version.

## What you get

Expand Down Expand Up @@ -155,7 +155,7 @@ with gconn.local_ray_context():

## 4. PyTorch DataLoader via the Permutation API

The training script doesn't know about the filter — it opens a view by name and reads through the [`Permutation`](/training/) API. Each DataLoader worker reopens its own connection lazily, reads Arrow batches directly from Lance (zero-copy, no intermediate file format), and the collate function decodes the whole batch in one pass. `Permutation` provides random-access indexing over the table, so shuffling is a cheap pointer rewrite rather than a full-dataset shuffle on disk.
The training script doesn't know about the filter — it opens a view by name and reads through the [`Permutation`](/training/data-loading) API. Each DataLoader worker reopens its own connection lazily, reads Arrow batches directly from Lance (zero-copy, no intermediate file format), and the collate function decodes the whole batch in one pass. `Permutation` provides random-access indexing over the table, so shuffling is a cheap pointer rewrite rather than a full-dataset shuffle on disk.

```py Python icon=Python
import lancedb
Expand Down
2 changes: 1 addition & 1 deletion docs/training/torch.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ for batch in dataloader:
```

Although the `Table` class in LanceDB implements the `torch.utils.data.Dataset` interface, you may find that using
a table [Permutation](/training/) is more flexible.
a table [Permutation](/training/data-loading) is more flexible.

```py Python icon=Python
from lancedb.permutation import Permutation
Expand Down
2 changes: 1 addition & 1 deletion docs/training/why-lancedb.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -103,7 +103,7 @@ access patterns in one format instead of scattering them across task-specific st
## Next steps

<CardGroup cols={2}>
<Card title="Data loading and shuffles" icon="boxes-stacked" href="/training/">
<Card title="Data loading and shuffles" icon="boxes-stacked" href="/training/data-loading">
Learn how to use LanceDB permutations to select rows, project columns, split datasets, and shuffle training reads.
</Card>
<Card title="PyTorch integration" icon="fire" href="/training/torch">
Expand Down
Loading