newsletter-august-2026
ChanChan Mao
data-mining-challenge-in-physical-ai
Lei Xu
feature-engineering-examples
Justin Miller
announcing-reverie-summit-2026
LanceDB
newsletter-july-2026
ChanChan Mao
crewai-rebuilt-agent-memory-on-lancedb
CrewAI
data-loading-guide
Weston Pace
one-table-to-train-your-robot-lancedb-as-the-data-layer-for-lerobot
Ayush Chaurasia
volcano-engine-lance-agent-memory
Bytedance
make-handwritten-notes-searchable-optimizing-an-ocr-pipeline-with-lancedb
Prashanth Rao
china-merchants-lancedb-story
China Merchants Lion Rock AI Lab
rabitq-gets-faster-higher-recall-lower-latency-query-time-control
Yang Cen
newsletter-june-2026
ChanChan Mao
from-messy-pdfs-to-verifiable-answers-with-liteparse-and-lancedb
Prashanth Rao
Clelia Astra Bertelli
faster-vlm-fine-tuning-with-materialized-model-features-in-lancedb
Prashanth Rao
Ayush Chaurasia
lance-blob-v2-late-materialization-for-large-binary-data-in-spark
Drew Gallardo
semantic-memory-for-hermes-agent-with-lancedb
Prashanth Rao
a-metadata-benchmark-of-lance-delta-lake-and-iceberg-on-s3
Jack Ye
scalable-feature-engineering-on-multimodal-datasets
Prashanth Rao
stable-worldmodel-a-high-performance-platform-for-reproducible-world-model-research
Ayush Chaurasia
Quentin Lhoest
Lucas Maes
Quentin Le Lidec
reproducible-data-curation-in-the-multimodal-lakehouse
Prashanth Rao
newsletter-may-2026
ChanChan Mao
newsletter-april-2026
ChanChan Mao
how-lancedb-accelerates-vector-search-at-10-billion-scale
Yang Cen
opensearch-vs-lancedb-for-vector-search-query-cost-and-infrastructure
Justin Miller
volcano-engine-autonomous-driving-data-lake-solution
Kejian Ju
unifying-the-av-ml-stack-lancedb
Ayush Chaurasia
lance-json-support-why-you-might-not-really-need-variant
Jack Ye
building-a-storage-format-for-the-next-era-of-biology
Pavan Ramkumar
newsletter-march-2026
ChanChan Mao
smart-parsing-meets-sharp-retrieval-combining-liteparse-and-lancedb
Clelia Astra Bertelli
Prashanth Rao
lance-format-v2-2-benchmarks-half-the-storage-none-of-the-slowdown
Xuanwo
make-your-sql-workflows-multimodal-with-lancedb-x-duckdb
Prashanth Rao
agentic-coding-as-community-stewardship
Xuanwo
what-we-mean-by-multimodal
Prashanth Rao
ai-native-development-local-continue-lancedb
Ty Dunn
lance-file-format-2-2-taming-complex-data
Xuanwo
lance-blob-v2
Xuanwo
Jack Ye
openclaw-lancedb-memory-layer
Xuanwo
Prashanth Rao
openclaw-lancedb-seed2
LanceDB
openclaw-memory-from-zero-to-lancedb-pro
Prashanth Rao
upload-lance-datasets-to-hf-hub
Prashanth Rao
zero-shot-image-classification-with-vector-search
Vipul Maheshwari
werides-data-platform-transformation-how-lancedb-fuels-model-development-velocity
Qian Zhu
Fei Chen
training-a-variational-autoencoder-from-scratch-with-the-lance-file-format
LanceDB
track-ai-trends-crewai-agents-rag
LanceDB
tokens-per-second-is-not-all-you-need
Mingran Wang
Tan Li
the-future-of-open-source-table-formats-iceberg-and-lance
Jack Ye
the-case-for-random-access-i-o
LanceDB
series-a-funding
Chang She
semanticdotart
Ayush Chaurasia
second-dinners-secret-weapon-lancedb-powered-rag-for-faster-smarter-game-development
Qian Zhu
search-within-an-image-331b54e4285e
Kaushal Choudhary
scalable-computer-vision-with-lancedb-voxel51-d8b65066d5f6
LanceDB
rethinking-table-file-paths-lance-multi-base-layout
Jack Ye
rag-isnt-one-size-fits-all
Leonard Marcq
python-package-to-convert-image-datasets-to-lance-type
Vipul Maheshwari
one-million-iops
Weston Pace
november-feature-roundup
Will Jones
newsletter-september-2025
Jasmine Wang
newsletter-october-2025
Jasmine Wang
newsletter-november-2025
ChanChan Mao
newsletter-june-2025
David Myriel
newsletter-july-2025
Jasmine Wang
newsletter-january-2026
ChanChan Mao
newsletter-february-2026
ChanChan Mao
newsletter-december-2025
ChanChan Mao
newsletter-august-2025
Jasmine Wang
my-summer-internship-experience-at-lancedb-2
Raunak Sinha
my-simd-is-faster-than-yours-fb2989bf25e7
LanceDB
multimodal-myntra-fashion-search-engine-using-lancedb
LanceDB
multimodal-lakehouse
David Myriel
multi-document-agentic-rag-a-walkthrough
Vipul Maheshwari
modified-rag-parent-document-bigger-chunk-retriever-62b3d1e79bc6
Mahesh Deshwal
memgpt-os-inspired-llms-that-manage-their-own-memory-793d6eed417e
Ayush Chaurasia
late-interaction-efficient-multi-modal-retrievers-need-more-than-just-a-vector-index
Ayush Chaurasia
lancedb-x-continue
LanceDB
lance-x-huggingface-a-new-era-of-sharing-multimodal-data
Prashanth Rao
Quentin Lhoest
Xuanwo
Ayush Chaurasia
lance-x-duckdb-sql-retrieval-on-the-multimodal-lakehouse-format
Xuanwo
lance-windows-windows-lance
Chang She
lance-v2
Weston Pace
lance-namespace-lancedb-and-ray
Jack Ye
lance-file-2-1-stable
Weston Pace
lance-file-2-1-smaller-and-simpler
Weston Pace
lance-data-viewer
Gordon Murray
lance-community-governance
Jack Ye
introducing-lance-namespace-spark-integration
Jack Ye
implementing-corrective-rag-in-the-easiest-way-2
LanceDB
hybrid-search-rag-for-real-life-production-grade-applications-e1e727b3965a
Mahesh Deshwal
hybrid-search-combining-bm25-and-semantic-search-for-better-results-with-lan-1358038fe7e6
LanceDB
hybrid-search-and-custom-reranking-with-lancedb-4c10a6a3447e
LanceDB
how-to-reduce-hallucinations-from-llm-powered-agents-using-long-term-memory-72f262c3cc1f
Tevin Wang
guide-to-use-contextual-retrieval-and-prompt-caching-with-lancedb
LanceDB
grpo-understanding-and-fine-tuning-the-next-gen-reasoning-model-2
Mahesh Deshwal
graphrag-hierarchical-approach-to-retrieval-augmented-generation
Akash Desai
gpu-accelerated-indexing-in-lancedb-27558fa7eee5
LanceDB
geo-support
Jack Ye
geneva-twelvelabs
David Myriel
geneva-feature-engineering
Jonathan Hsieh
from-bi-to-ai-lance-and-iceberg
Jack Ye
Prashanth Rao
fluss-integration
Wayne Wang
file-readers-in-depth-parallelism-without-row-groups
Weston Pace
feature-rabitq-quantization
David Myriel
Yang Cen
feature-full-text-search
David Myriel
enhance-rag-integrate-contextual-compression-and-filtering-for-precision-a29d4a810301
Kaushal Choudhary
effortlessly-loading-and-processing-images-with-lance-a-code-walkthrough
LanceDB
designing-a-table-format-for-ml-workloads
Weston Pace
custom-dataset-for-llm-training-using-lance
LanceDB
creating-a-fintech-agent
Vipul Maheshwari
convert-any-image-dataset-to-lance
LanceDB
columnar-file-readers-in-depth-structural-encoding
Weston Pace
columnar-file-readers-in-depth-repetition-definition-levels
Weston Pace
columnar-file-readers-in-depth-compression-transparency
Weston Pace
columnar-file-readers-in-depth-column-shredding
Weston Pace
columnar-file-readers-in-depth-backpressure
Weston Pace
columnar-file-readers-in-depth-apis-and-fusion
Weston Pace
chunking-techniques-with-langchain-and-llamaindex
Prashant Kumar
chunking-analysis-which-is-the-right-chunking-approach-for-your-language
Shresth Shukla
chat-with-csv-excel-using-lancedb
LanceDB
case-study-netflix
David Myriel
case-study-dosu
Qian Zhu
Michael Ludden
case-study-cognee
David Myriel
Vasilije Markovic
case-study-coderabbit
Qian Zhu
building-rag-on-codebases-part-2
Sankalp Shubham
building-rag-on-codebases-part-1
Sankalp Shubham
branching-and-shallow-clone
Jack Ye
better-rag-with-active-retrieval-augmented-generation-flare-3b66646e2a9f
LanceDB
benchmarking-random-access-in-lance
Chang She
benchmarking-lancedb-92b01032874a-2
LanceDB
benchmarking-cohere-reranker-with-lancedb
LanceDB
anythingllms-competitive-edge-lancedb-for-seamless-rag-and-agent-workflows
Ayush Chaurasia
announcing-lance-sdk
Weston Pace
agentic-rag-using-langgraph-building-a-simple-customer-support-autonomous-agent
LanceDB
advanced-rag-precise-zero-shot-dense-retrieval-with-hyde-0946c54dfdcb
LanceDB
accelerate-vector-search-applications-using-openvino-lancedb
LanceDB
a-primer-on-text-chunking-and-its-types-a420efc96a13
Prashant Kumar
a-practical-guide-to-training-custom-rerankers
Ayush Chaurasia
a-practical-guide-to-fine-tuning-embedding-models
Ayush Chaurasia
keep-your-data-fresh-with-cocoindex-and-lancedb
Prashanth Rao
Linghua Jin

🎤 Reverie, Nov 5, ⚡ ML Data Loading Performance Guide, 🧠 CrewAI Cognitive Memory on LanceDB

September 9, 2026
Newsletter

🎤 Reverie, the Summit for AI Builders

Every generative model is a kind of dream, a plausible world rendered from what it has seen. Reverie, the summit for AI builders by LanceDB, is taking place on November 5 in San Francisco. Researchers and engineers behind frontier models, world models, generative video, and physical AI go deep on the data systems those dreams are made on.

Featuring speakers from NVIDIA, Runway, Luma, Applied Intuition, Cruise, Adobe, Exa, and more.

Apply to Attend → | Read more →

🧠 Why CrewAI Rebuilt Agent Memory on LanceDB, Powering 2B+ Agent Executions

CrewAI replaced a two-system memory stack with a single LanceDB table storing vectors, metadata, and multimodal data together, scoring recall by similarity, recency, and importance so old critical decisions still surface over trivial recent ones. Now CrewAI's default vector backend, it powers 12M monthly downloads and 2B+ agent executions.

Read more →

⚡ Data Loading for AI/ML: A Comprehensive Guide

GPU training throughput rarely bottlenecks on I/O — the real trap is the CPU stage, where JPEG decoding and tokenization can starve even a capable pipeline. Weston Pace walks through the three-stage model with concrete numbers (16K tokens/sec on an L40S, 1GB/s per core) and practical PyTorch + LanceDB guidance.

Read more →

📚 Also Published

📸 Actuate & Ray Summit in SF

LanceDB team at Actuate and Ray Summit this month! Thanks to everyone who stopped by our booth and said hi. We really enjoyed meeting so many of you in person and talking about data infrastructure for physical AI, robotics, and foundation models. See you at our own Reverie on November 5 👋

📅 Upcoming Events

Reverie — November 5, 2026 · San Francisco, CA

Reverie is a one-day technical summit for researchers and engineers building the data systems behind world models, generative video, physical AI, multimodal search, and agentic research.

Apply to Attend →

🏗️ LanceDB Enterprise Updates

Performance

  • Less job-history scanning on cleanup — Per-job scratch storage removes a cleanup job's need to scan every other job's record for what's expired; on a 604-record benchmark, that scan alone cost ~600 sequential storage reads and 15.2 of 26.3 total seconds (58%).
  • Fewer redundant manifest reads on WAL nodes — Reusing a shard's own manifest cache instead of rebuilding it each read cuts a flush operation's storage requests from 132 to 12, and stops idle tables from burning ~17 requests/minute on background GC and compaction checks with no client traffic.
  • Concurrent SSTable fetch in WAL compaction — Fetching a compaction pass's SSTables concurrently instead of one at a time cuts fetch time by roughly 4x at the default fan-out: an 8-SSTable pass drops from 1.6s to 0.4s, and a 32-SSTable pass from 6.4s to 1.6s.
  • Bounded admin traffic on large clusters — Capping concurrency on cluster-status requests and polling only the replica-group members a job actually targets, instead of every member cluster-wide, prevents a connection spike on the query node as replica-group counts grow.
  • Streaming computed-column refresh — Streaming a computed column's output directly into its fragment file, instead of buffering the whole output twice, removes the artificial cap that previously bounded how much a single refresh batch could compute at once.

Features

Feature Description
Materialized views SQL now supports CREATE, REFRESH, and SHOW MATERIALIZED VIEW, a filtered or projected view of another table that refreshes as a background job, with a validation token guarding against refreshing a view dropped and recreated under the same name.
SQL-defined computed columns SQL now supports registering Python functions as computed columns, with WHERE-filtered refresh to recompute only matching rows and GPU-backed execution for functions that need it, bringing Feature Engineering's computed-column workflow into SQL.
Table cloning in SQL CREATE TABLE ... CLONE now works from the SQL shell and Flight SQL, closing the last surface that lacked it: clone a table as of its latest version, a specific version, a tag, or a timestamp, without copying the underlying data.
Column lineage graph A new lineage view shows which upstream columns and Feature Engineering jobs produced each column, rendered as an interactive graph with search and click-to-pin, keeping a complex table's derivation history legible instead of an unreadable tangle.
Row and fragment sampling The SQL:2003 TABLESAMPLE clause is now supported for both row-level and whole-fragment sampling, by percentage or exact row count, with a repeatable seed for reproducible samples and support for sampling one side of a join independently.

🌟 Open Source Releases

Project Description
Lance v3.0.2 – v11.0.0
Release notes
• ACORN-1 traversal cuts worst-case HNSW search latency 10–250x when a filter clusters away from the query's region, at up to 4 points of recall (opt-in via ApproxMode::Fast) (#7927); a new MAXSCORE algorithm for pure-SHOULD FTS queries cuts candidate probes up to 53.6x in benchmark testing (#8474); and FTS scorers can now compose across columns in one shared row-address domain (#8685)
• Segmented indexes now support ngram (#7244), bloom filters (#7925), and RTree (#7932); zone maps extended to all data types (#8017) with seeds written into data file footers (#7427)
• New Dataset.migrate_to_stable_row_ids migration method (#8521), cross-store deep_clone support (#7545), compaction limits via max_source_rows/max_source_bytes (#8235), and per-fragment column writes that survive compaction (#8313)
LanceDB v0.37.1 – v0.38.0
Release notes
• Materialized views land in the client SDKs — declare and refresh on local tables (#3930, #4010) with Python (#3933) and Node (#3935) bindings, alongside SQL-declared computed columns (#3937) with async refresh returning a job handle (#3939)
• New Function framework for registering scalar UDFs with typed wire contracts (#3985), grouped column bindings (#3994), GPU resource requirements (#4085), and Blob v2 signatures (#4091)
• Streaming dataset adds backpressure on its post-transform queue (#3897) and sequence packing (#3920); the data loader now reads remote tables (#3981) pinned to a fixed base table version (#3982)
lance-namespace v0.11.0 – v0.11.1
Release notes
InsertIntoTableResponse now returns num_inserted_rows and version fields (#359)
• Support for declaring and backfilling expression-computed columns (#360)
• Vector index build parameters (e.g. IVF/PQ settings) can now be specified via CreateTableIndexRequest (#361)
lance-trino v0.3.4
Release notes
• Fixed Substrait name emission for structs nested inside lists, improving compatibility with complex schema queries (#227)
• Optimized COUNT() queries over non-null constants (#155)

🫶 Community Contributions

Thank you to contributors from Netflix, Bytedance, Adobe, Intel, Microsoft, Pinterest, Huawei for improvements across storage, indexing, query execution, distributed processing, and ecosystem integrations in LanceDB, Lance, and the broader ecosystem.

Notable contributions this month:

  • @yanghua — Added a pluggable cache backend API to the Java SDK, letting custom caching strategies register and switch at runtime on top of Lance's existing cache registry
  • @zhangyue19921010 — Introduced compaction controls with max_source_rows and max_source_bytes limits, enabling fine-grained resource management during file compaction
  • @sezruby — Exposed ArrowArrayStream export on LanceScanner in Java, enabling zero-copy interoperability with Arrow-native tooling
  • @morales-t-netflix — Added per-value support for fixed-length packed structs, enabling efficient columnar encoding for complex nested types
  • @leohoare — Implemented ACORN-1 traversal for prefiltered HNSW search, delivering 10–250x faster filtered vector queries at a small (≤4 point) recall tradeoff
  • @xtangxtang — Fixed fp16 IVF partition assignment to route through AMX-FP16 GEMM, recovering misassigned vectors and accelerating vector indexing on Intel hardware
  • @professor-moody — Contributed critical security fix bounds-checking length prefix in VariableFullZipDecoder, preventing potential memory safety issues
  • @XuQianJin-Stars — Implemented safe commit with ConditionalPutCommitHandler for GooseFS, enabling consistent writes on distributed filesystems
  • @wombatu-kun — Fixed stale per-segment rows in vector search after in-place column value updates, keeping the index consistent with the latest data

A heartfelt thank you to our community contributors of Lance and LanceDB this past month:

@1fanwang@3286360470@a-erofeev@adityaj0@ali2arslan@anirudh-s-kumar@anonx3247@antio2@ar-maan05@arielyes@atirna@beinan@brunosrz@clearlove10-c@com-junkawasaki@cswpy@dawid0309@dcfocus@ddupg@dentiny@divyanshus2404@dshepelev15@dtolnay@dubin555@ecthlion@everysympathy@fangbo@fanng1@farazshaikh@fzowl@geserdugarov@haochengliu@hellower@hfutatzhanghb@huahuay@igorganapolsky@ilya-zlobintsev@isaac-dasari@ivscheianu@j7nhai@jackylee-ch@jagrutipatilp@jay-ju@jayson-huang@jerryjch@jiaoew1991@jiaqizho@jimmy-xie-fleet@jo-migo@jonasdedden@jsnider3@julianyg@kamronis@kangnan-li@keunhong@leoreeyang@lh-kevin@lichuang@luciferyang@lucyge2022@majin1102@mannxo@maswin@medisean@mikemikimike@mikewhb@mmatczuk@nyl3532016@paramt@pengw0048@pjdurden@primorlee@prrao87@puchengy@qiuyuhang@ragingkore@ragnorc@raygao25@ringfa11@risto0211@roridemonslayer@rupertmaiti2005@sbrunk@seven7763@sliortega295-ops@sravan1011@stevestevenpoor@tandede@touch-of-grey@u70b3@valkum@vatharevinayak@vinaysurtani@vip892766gma@weimingdiit@winklemad@wirybeaver@xiaguanglei@xixigoodluck@xloya@yentur@yichenw@yuvalif@yuw1@zackfairts@zhangstar333@zouhuajian@zyt-yt

🤝 Lance Community Sync Recap

The community syncs this month covered the Lance 10.0.0 SDK release, which included a critical encoding fix backported to all 0.x versions, along with a refactor of file version handling. Discussion also focused on the Lance 11.0.0 SDK beta development and its progression to RC2, with work done to address benchmark regressions before release. Process changes were announced for format-change votes, moving from GitHub discussions to pull requests with a shortened 72-hour window. The team also discussed proposals for experimental feature flags and native video/blob data types in Lance.

The next Lance Community Sync will take place on Thursday, September 10 @ 9am PT.

ChanChan Mao
Developer Relations @ LanceDB

🎤 Reverie, Nov 5, ⚡ ML Data Loading Performance Guide, 🧠 CrewAI Cognitive Memory on LanceDB

ChanChan Mao
September 8, 2026
newsletter-august-2026

Data Loading for AI/ML: A Comprehensive Guide

Weston Pace
July 22, 2026
data-loading-guide

Announcing Reverie Summit: What AI’s Next Breakthroughs Are Made On

LanceDB
August 3, 2026
announcing-reverie-summit-2026