data-mining-challenge-in-physical-ai
Lei Xu
feature-engineering-examples
Justin Miller
announcing-reverie-summit-2026
LanceDB
newsletter-july-2026
ChanChan Mao
crewai-rebuilt-agent-memory-on-lancedb
CrewAI
data-loading-guide
Weston Pace
one-table-to-train-your-robot-lancedb-as-the-data-layer-for-lerobot
Ayush Chaurasia
volcano-engine-lance-agent-memory
Bytedance
make-handwritten-notes-searchable-optimizing-an-ocr-pipeline-with-lancedb
Prashanth Rao
china-merchants-lancedb-story
China Merchants Lion Rock AI Lab
rabitq-gets-faster-higher-recall-lower-latency-query-time-control
Yang Cen
newsletter-june-2026
ChanChan Mao
from-messy-pdfs-to-verifiable-answers-with-liteparse-and-lancedb
Prashanth Rao
Clelia Astra Bertelli
faster-vlm-fine-tuning-with-materialized-model-features-in-lancedb
Prashanth Rao
Ayush Chaurasia
lance-blob-v2-late-materialization-for-large-binary-data-in-spark
Drew Gallardo
semantic-memory-for-hermes-agent-with-lancedb
Prashanth Rao
a-metadata-benchmark-of-lance-delta-lake-and-iceberg-on-s3
Jack Ye
scalable-feature-engineering-on-multimodal-datasets
Prashanth Rao
stable-worldmodel-a-high-performance-platform-for-reproducible-world-model-research
Ayush Chaurasia
Quentin Lhoest
Lucas Maes
Quentin Le Lidec
reproducible-data-curation-in-the-multimodal-lakehouse
Prashanth Rao
newsletter-may-2026
ChanChan Mao
newsletter-april-2026
ChanChan Mao
how-lancedb-accelerates-vector-search-at-10-billion-scale
Yang Cen
opensearch-vs-lancedb-for-vector-search-query-cost-and-infrastructure
Justin Miller
volcano-engine-autonomous-driving-data-lake-solution
Kejian Ju
unifying-the-av-ml-stack-lancedb
Ayush Chaurasia
lance-json-support-why-you-might-not-really-need-variant
Jack Ye
building-a-storage-format-for-the-next-era-of-biology
Pavan Ramkumar
newsletter-march-2026
ChanChan Mao
smart-parsing-meets-sharp-retrieval-combining-liteparse-and-lancedb
Clelia Astra Bertelli
Prashanth Rao
lance-format-v2-2-benchmarks-half-the-storage-none-of-the-slowdown
Xuanwo
make-your-sql-workflows-multimodal-with-lancedb-x-duckdb
Prashanth Rao
agentic-coding-as-community-stewardship
Xuanwo
what-we-mean-by-multimodal
Prashanth Rao
ai-native-development-local-continue-lancedb
Ty Dunn
lance-file-format-2-2-taming-complex-data
Xuanwo
lance-blob-v2
Xuanwo
Jack Ye
openclaw-lancedb-memory-layer
Xuanwo
Prashanth Rao
openclaw-lancedb-seed2
LanceDB
openclaw-memory-from-zero-to-lancedb-pro
Prashanth Rao
upload-lance-datasets-to-hf-hub
Prashanth Rao
zero-shot-image-classification-with-vector-search
Vipul Maheshwari
werides-data-platform-transformation-how-lancedb-fuels-model-development-velocity
Qian Zhu
Fei Chen
training-a-variational-autoencoder-from-scratch-with-the-lance-file-format
LanceDB
track-ai-trends-crewai-agents-rag
LanceDB
tokens-per-second-is-not-all-you-need
Mingran Wang
Tan Li
the-future-of-open-source-table-formats-iceberg-and-lance
Jack Ye
the-case-for-random-access-i-o
LanceDB
series-a-funding
Chang She
semanticdotart
Ayush Chaurasia
second-dinners-secret-weapon-lancedb-powered-rag-for-faster-smarter-game-development
Qian Zhu
search-within-an-image-331b54e4285e
Kaushal Choudhary
scalable-computer-vision-with-lancedb-voxel51-d8b65066d5f6
LanceDB
rethinking-table-file-paths-lance-multi-base-layout
Jack Ye
rag-isnt-one-size-fits-all
Leonard Marcq
python-package-to-convert-image-datasets-to-lance-type
Vipul Maheshwari
one-million-iops
Weston Pace
november-feature-roundup
Will Jones
newsletter-september-2025
Jasmine Wang
newsletter-october-2025
Jasmine Wang
newsletter-november-2025
ChanChan Mao
newsletter-june-2025
David Myriel
newsletter-july-2025
Jasmine Wang
newsletter-january-2026
ChanChan Mao
newsletter-february-2026
ChanChan Mao
newsletter-december-2025
ChanChan Mao
newsletter-august-2025
Jasmine Wang
my-summer-internship-experience-at-lancedb-2
Raunak Sinha
my-simd-is-faster-than-yours-fb2989bf25e7
LanceDB
multimodal-myntra-fashion-search-engine-using-lancedb
LanceDB
multimodal-lakehouse
David Myriel
multi-document-agentic-rag-a-walkthrough
Vipul Maheshwari
modified-rag-parent-document-bigger-chunk-retriever-62b3d1e79bc6
Mahesh Deshwal
memgpt-os-inspired-llms-that-manage-their-own-memory-793d6eed417e
Ayush Chaurasia
late-interaction-efficient-multi-modal-retrievers-need-more-than-just-a-vector-index
Ayush Chaurasia
lancedb-x-continue
LanceDB
lance-x-huggingface-a-new-era-of-sharing-multimodal-data
Prashanth Rao
Quentin Lhoest
Xuanwo
Ayush Chaurasia
lance-x-duckdb-sql-retrieval-on-the-multimodal-lakehouse-format
Xuanwo
lance-windows-windows-lance
Chang She
lance-v2
Weston Pace
lance-namespace-lancedb-and-ray
Jack Ye
lance-file-2-1-stable
Weston Pace
lance-file-2-1-smaller-and-simpler
Weston Pace
lance-data-viewer
Gordon Murray
lance-community-governance
Jack Ye
introducing-lance-namespace-spark-integration
Jack Ye
implementing-corrective-rag-in-the-easiest-way-2
LanceDB
hybrid-search-rag-for-real-life-production-grade-applications-e1e727b3965a
Mahesh Deshwal
hybrid-search-combining-bm25-and-semantic-search-for-better-results-with-lan-1358038fe7e6
LanceDB
hybrid-search-and-custom-reranking-with-lancedb-4c10a6a3447e
LanceDB
how-to-reduce-hallucinations-from-llm-powered-agents-using-long-term-memory-72f262c3cc1f
Tevin Wang
guide-to-use-contextual-retrieval-and-prompt-caching-with-lancedb
LanceDB
grpo-understanding-and-fine-tuning-the-next-gen-reasoning-model-2
Mahesh Deshwal
graphrag-hierarchical-approach-to-retrieval-augmented-generation
Akash Desai
gpu-accelerated-indexing-in-lancedb-27558fa7eee5
LanceDB
geo-support
Jack Ye
geneva-twelvelabs
David Myriel
geneva-feature-engineering
Jonathan Hsieh
from-bi-to-ai-lance-and-iceberg
Jack Ye
Prashanth Rao
fluss-integration
Wayne Wang
file-readers-in-depth-parallelism-without-row-groups
Weston Pace
feature-rabitq-quantization
David Myriel
Yang Cen
feature-full-text-search
David Myriel
enhance-rag-integrate-contextual-compression-and-filtering-for-precision-a29d4a810301
Kaushal Choudhary
effortlessly-loading-and-processing-images-with-lance-a-code-walkthrough
LanceDB
designing-a-table-format-for-ml-workloads
Weston Pace
custom-dataset-for-llm-training-using-lance
LanceDB
creating-a-fintech-agent
Vipul Maheshwari
convert-any-image-dataset-to-lance
LanceDB
columnar-file-readers-in-depth-structural-encoding
Weston Pace
columnar-file-readers-in-depth-repetition-definition-levels
Weston Pace
columnar-file-readers-in-depth-compression-transparency
Weston Pace
columnar-file-readers-in-depth-column-shredding
Weston Pace
columnar-file-readers-in-depth-backpressure
Weston Pace
columnar-file-readers-in-depth-apis-and-fusion
Weston Pace
chunking-techniques-with-langchain-and-llamaindex
Prashant Kumar
chunking-analysis-which-is-the-right-chunking-approach-for-your-language
Shresth Shukla
chat-with-csv-excel-using-lancedb
LanceDB
case-study-netflix
David Myriel
case-study-dosu
Qian Zhu
Michael Ludden
case-study-cognee
David Myriel
Vasilije Markovic
case-study-coderabbit
Qian Zhu
building-rag-on-codebases-part-2
Sankalp Shubham
building-rag-on-codebases-part-1
Sankalp Shubham
branching-and-shallow-clone
Jack Ye
better-rag-with-active-retrieval-augmented-generation-flare-3b66646e2a9f
LanceDB
benchmarking-random-access-in-lance
Chang She
benchmarking-lancedb-92b01032874a-2
LanceDB
benchmarking-cohere-reranker-with-lancedb
LanceDB
anythingllms-competitive-edge-lancedb-for-seamless-rag-and-agent-workflows
Ayush Chaurasia
announcing-lance-sdk
Weston Pace
agentic-rag-using-langgraph-building-a-simple-customer-support-autonomous-agent
LanceDB
advanced-rag-precise-zero-shot-dense-retrieval-with-hyde-0946c54dfdcb
LanceDB
accelerate-vector-search-applications-using-openvino-lancedb
LanceDB
a-primer-on-text-chunking-and-its-types-a420efc96a13
Prashant Kumar
a-practical-guide-to-training-custom-rerankers
Ayush Chaurasia
a-practical-guide-to-fine-tuning-embedding-models
Ayush Chaurasia
keep-your-data-fresh-with-cocoindex-and-lancedb
Prashanth Rao
Linghua Jin

Turning Fleet Data Into Better Models: The Data Mining Challenge in Physical AI

September 3, 2026
Physical AI

Physical AI systems generate an enormous amount of data.

Every vehicle, robot, and sensor-equipped machine is continuously producing camera frames, LiDAR, telemetry, control signals, trajectories, model outputs, annotations, and other state. At any meaningful fleet size, collecting data is not really the hard problem anymore.

The hard problem is finding the small fraction of that data that can make the model better.

A robot fails to grasp an object under unusual lighting. An autonomous vehicle brakes harder than expected at an intersection. A perception model behaves strangely around a particular combination of rain, pedestrians, and road geometry.

You may have already captured hundreds of similar examples. But if they are buried inside petabytes of fleet logs, they are not very useful.

This is why I think data mining is becoming one of the most important infrastructure problems in physical AI. The data system needs to turn raw fleet experience into searchable, reproducible datasets for training and evaluation.

The better your model gets, the harder the data problem becomes

Early in model development, almost any additional data can help.

That changes as the model improves.

Most of the fleet eventually represents situations the model already handles well. Driving straight on a clear road for the ten-millionth time probably has much less training value than the first examples did.

The useful examples move further into the long tail.

You start asking questions like:

  • Show me every gripper slip where force variance increased immediately before failure.
  • Find hard decelerations in Boston.
  • Find left turns between 6 PM and 8 PM in this geofence where two or three pedestrians crossed from the left side of the vehicle.
  • Find clips visually similar to this failure, even if nobody previously labeled them.

The difficulty is that the rarer the event becomes, the more data you have to search to find it. This is a slightly unusual infrastructure problem. A two-second query instead of a 200 ms query might not matter very much to the researcher. Being able to search 100 billion examples instead of 100 million matters enormously.

That distinction comes up repeatedly when we talk with robotics and autonomous-vehicle teams. The challenge is not just low-latency retrieval. It is making very large multimodal datasets searchable enough to mine the long tail. This is fundamentally a scale problem where supporting tens or hundreds of billions of rows matters much more than shaving small amounts of interactive latency.

As models mature, useful training examples become a smaller fraction of the fleet. Data mining has to search more data to find fewer high-value examples.

Data mining is the center of the Physical AI data flywheel

A new failure happens in the field. You find similar examples across the historical fleet. You enrich and curate those examples into a dataset. You train or evaluate a new model. You deploy it. Then you mine the next generation of failures.

The faster this loop runs, the faster real-world experience turns into model quality.

But there is an important detail here. Mining does not simply mean writing SQL over a table.

Physical AI data starts as video, images, point clouds, ROS bags, MCAP files, trajectories, sensor streams, and other complex structures. The thing you want to search for often does not exist as a column yet. You have to create it.

A researcher might be mining grasp failures and realize the useful signal is not an existing label but a pattern in joint position, velocity, force, and torque over the two seconds before failure. An AV researcher might define a new metric from ego kinematics, object tracks, and planner outputs to capture a long-tail interaction that existing metrics miss. Or a team might run a video-language model over stored clips to generate scene descriptions, direct video embeddings, or model-error features, then use those outputs to retrieve similar failures. In practice, the features needed for mining are often discovered during mining itself, which makes feature engineering part of the loop.

Physical AI improves through a closed-loop data engine. Every deployment produces new experience, which becomes the input for the next mining and training cycle.

What this workflow looks like today

In practice, a lot of Physical AI stacks evolved one workload at a time.

The raw data lands in object storage. ROS bags or MCAP files are decoded with Spark, Ray, Dataflow, or custom pipelines. Some subset is loaded into a visualization system such as Foxglove or Rerun. Metadata goes into a warehouse or analytical database. Embeddings go into a vector database or search engine. Training data gets rewritten again into WebDataset, LeRobot, or another training-optimized representation.

Then somebody builds an API that tries to make all of those systems look like one thing.

The problem is that you have to synchronize copies. You have to maintain lineage between the raw event, its derived features, the search index, and the eventual training example. Adding a feature can mean another large backfill and sometimes another physical representation of the dataset.

This gets particularly ugly with multimodal data because the feature computation itself can be substantial. A new column might require downloading an MP4, extracting a time window, decoding frames, converting them into tensors, running a PyTorch model, and writing an embedding back into the dataset.

Physical AI teams often maintain several representations of the same underlying experience for playback, analytics, search, feature engineering, and training.

Feature engineering changes what is searchable

This is the part of data mining that I think is sometimes underestimated.

Imagine that you want every left turn between 6 PM and 8 PM inside this geofence where two or three pedestrians cross from the left side of the vehicle.

Time is structured metadata. Location can be indexed geospatially. But “left turn” may need to be derived from trajectory data. “Two or three pedestrians” may require a perception model. “Crossing from the left side” may require temporal reasoning over a video clip.

The researcher cannot query these concepts until the system turns them into something computationally searchable. For example, leveraging yaw rate and speed, add a column is_left_turn to find left turns in the data set:

@geneva.udf(
    data_type=pa.bool_(),
    input_columns=[
        "telemetry_time_s",
        "yaw_rate_down_rps",
        "steering_angle_deg",
        "car_speed_mps",
    ],
    num_cpus=1.0,
    num_gpus=0.0,
    version="left-turn-v1",
)
def detect_left_turn(
    telemetry_time_s: list[float],
    yaw_rate_down_rps: list[float],
    steering_angle_deg: list[float],
    car_speed_mps: list[float],
) -> bool:
    import numpy as np

    t = np.asarray(telemetry_time_s, dtype=np.float64)
    yaw = np.asarray(yaw_rate_down_rps, dtype=np.float64)
    steering = np.asarray(steering_angle_deg, dtype=np.float64)
    speed = np.asarray(car_speed_mps, dtype=np.float64)
    n = min(len(t), len(yaw), len(steering), len(speed))
    if n < 2:
        return False
    t, yaw, steering, speed = t[:n], yaw[:n], steering[:n], speed[:n]

    valid = np.isfinite(t) & np.isfinite(yaw) & np.isfinite(steering) & np.isfinite(speed)
    t, yaw, steering, speed = t[valid], yaw[valid], steering[valid], speed[valid]
    if len(t) < 2:
        return False

    moving = speed >= 2.0
    candidate = moving & (yaw <= -0.08)  # negative down-axis yaw is left

    # Learn whether steering sign agrees or disagrees with yaw sign.
    informative = moving & (np.abs(yaw) >= 0.03) & (np.abs(steering) >= 1.0)
    if informative.sum() >= 20:
        corr = np.corrcoef(yaw[informative], steering[informative])[0, 1]
        if np.isfinite(corr) and abs(corr) >= 0.20:
            steering_in_yaw_sign = steering * (1.0 if corr > 0 else -1.0)
            candidate &= steering_in_yaw_sign <= -5.0

    typical_dt = float(np.median(np.diff(t)))
    max_gap = max(0.20, 2.5 * typical_dt)
    run_start = None
    previous_t = None
    for timestamp, is_candidate in zip(t, candidate, strict=True):
        if is_candidate and (previous_t is None or timestamp - previous_t <= max_gap):
            run_start = timestamp if run_start is None else run_start
            if timestamp - run_start >= 0.75:
                return True
        elif is_candidate:
            run_start = timestamp
        else:
            run_start = None
        previous_t = timestamp
    return False


detect_left_turn

if "is_left_turn" not in table.schema.names:
    table.add_columns({"is_left_turn": detect_left_turn})

with db.local_ray_context():
    left_turn_job = table.backfill(
        "is_left_turn",
        concurrency=1,
        task_size=1,
        max_checkpoint_size=1,
    )

table.checkout_latest()
table.search().select(["segment_id", "duration_s", "is_left_turn"]).to_pandas()

And we can leverage a pretrained, COCO-trained Faster R-CNN MobileNet detector once per worker and reuse it across rows. COCO's `person` class can be used to find pedestrians.

@geneva.udf(
    data_type=pa.int32(),
    input_columns=["video_hevc"],
    num_cpus=2.0,
    num_gpus=0.0,
    memory=2_000_000_000,
    version="pedestrian-count-v3-resnet50",
)
class PedestrianCounter(Callable):
    def __init__(
        self,
        score_threshold: float = 0.60,
        frame_stride: int = 50,
        max_frames: int = 12,
        inference_batch_size: int = 2,
    ) -> None:
        self.score_threshold = score_threshold
        self.frame_stride = frame_stride
        self.max_frames = max_frames
        self.inference_batch_size = inference_batch_size
        self._loaded = False

    def setup(self) -> None:
        import torch
        from torchvision.models.detection import (
            FasterRCNN_ResNet50_FPN_V2_Weights,
            fasterrcnn_resnet50_fpn_v2,
        )

        weights = FasterRCNN_ResNet50_FPN_V2_Weights.DEFAULT
        self.device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
        self.model = fasterrcnn_resnet50_fpn_v2(weights=weights)
        self.model.to(self.device).eval()
        self._loaded = True

    def _max_people_in_batch(self, images: list) -> int:
        import torch

        if not images:
            return 0
        with torch.inference_mode():
            predictions = self.model([image.to(self.device) for image in images])
        return max(
            int(((pred["labels"] == 1) & (pred["scores"] >= self.score_threshold)).sum().item())
            for pred in predictions
        )

    def __call__(self, video_hevc: bytes) -> int:
        import io

        import av
        from torchvision.transforms.functional import pil_to_tensor

        if not self._loaded:
            self.setup()

        scene_max = 0
        batch = []
        sampled = 0
        # Auto-detect raw HEVC (comma2k19) or MP4 (JAAD).
        with av.open(io.BytesIO(video_hevc)) as container:
            for frame_index, frame in enumerate(container.decode(video=0)):
                if frame_index % self.frame_stride != 0:
                    continue
                tensor = pil_to_tensor(frame.to_image()).float().div_(255.0)
                batch.append(tensor)
                sampled += 1
                if len(batch) == self.inference_batch_size:
                    scene_max = max(scene_max, self._max_people_in_batch(batch))
                    batch.clear()
                if sampled >= self.max_frames:
                    break
        if batch:
            scene_max = max(scene_max, self._max_people_in_batch(batch))
        return int(scene_max)


pedestrian_counter = PedestrianCounter()
pedestrian_counter

if "pedestrian_count" not in table.schema.names:
    table.add_columns({"pedestrian_count": pedestrian_counter})

with db.local_ray_context():
    pedestrian_job = table.backfill(
        "pedestrian_count",
        concurrency=1,
        task_size=1,
        max_checkpoint_size=1,
    )

table.checkout_latest()

Then we can search for scenes similar to a reference clip while filtering on the structured features we just created:

results = (
    table.search(
        example_clip_embedding,
        vector_column_name="left_camera_embedding",
    )
    .where("""
        city = 'boston'
        AND hour BETWEEN 18 AND 20
        AND is_left_turn = true
        AND pedestrian_count BETWEEN 2 AND 3
    """)
    .limit(100)
    .to_pandas()
)

This is why treating feature engineering, data management, and retrieval as completely separate systems becomes increasingly awkward. Mining is iterative. Researchers discover that they need a new signal, compute it, inspect the results, change it, backfill history, and query again. Raw sensor logs become useful for mining only after they are transformed into discrete attributes, booleans, numerical values, or semantic embeddings.

A useful data system should keep the raw experience and the searchable representation together

This is the architectural idea behind Lance and LanceDB. Instead of treating the raw multimodal data, analytical representation, embeddings, and training dataset as unrelated copies, we want them to behave like different views of the same underlying dataset.

The raw media and sensor data stay addressable. New features and model outputs can be added over time. The same dataset can support structured filters, full-text search, vector search, and other retrieval patterns. Historical data can be backfilled when researchers define a new feature.

And importantly, the output of mining does not have to become another disconnected data silo. It can become a versioned dataset or materialized subset that is used directly for training and evaluation.

LanceDB's direction here is to combine the pieces required for this loop, with multimodal storage, iterative feature engineering, historical backfills, indexing, hybrid retrieval, dataset creation, and training access over the same underlying data. The original data-mining design specifically calls out feature enrichment, automatic backfilling, hybrid search across SQL/full-text/vector/geospatial signals, and mining directly against source data.

A unified multimodal data layer lets teams use the same underlying dataset across playback, mining, feature engineering, training, and evaluation.

Mining is also how you build better evaluation sets

There is another benefit to this architecture. Every new failure mode gives you a candidate regression test. If a vehicle encounters a new scenario, the team can mine historically similar incidents and turn them into a versioned evaluation set. The next model should not only solve the new failure. It should continue to solve the previous ones.

Over time, the fleet itself becomes the source of increasingly difficult “golden” evaluation sets. This is especially useful in safety-critical systems. When something unusual occurs, engineers need to answer two questions quickly.

This is fundamentally a data retrieval and dataset-management problem. The original workflow similarly connects incident investigation with building regression sets from real-world failures.

A field failure should become a reusable regression test for every model that follows.

The data advantage in Physical AI will come from using experience better

The companies that win in Physical AI will not just be the ones that collect the most data. They will be the ones that can turn fleet experience into better training data faster by finding rare failures, computing new signals over historical data, and turning those results into reproducible datasets for training and evaluation.

As models improve, the valuable examples become harder to find, and the data system becomes part of the model-development loop itself. Data mining is what connects what happened in the real world to what the model learns next, and the teams that can run that loop efficiently will improve faster.

Lei Xu
Co-founder and CTO

Data Loading for AI/ML: A Comprehensive Guide

Weston Pace
July 22, 2026
data-loading-guide

Announcing Reverie Summit: What AI’s Next Breakthroughs Are Made On

LanceDB
August 3, 2026
announcing-reverie-summit-2026

⚡ Multi-Bit RaBitQ Without Refine, 🌋 Bytedance’s Lance-Based AI Stack, 🤖 Lance for Embodied AI Data

ChanChan Mao
July 31, 2026
newsletter-july-2026