Latest

6/recent/ticker-posts

Header Ads Widget

Inside Spark 4.2 ✨, Data Modeling Still Matters 📐, Faster Spark With Rust 🦀

DataFusion Comet speeds Spark reads on Iceberg by keeping Java planning while using Iceberg Rust for native Arrow execution. On a TPC-DS benchmark ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌  ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ ‌ 

TLDR

Together With String

TLDR Data 2026-07-20

Your agent is failing to fetch web data; the average failure rate for hard-target web fetches is 40%. Ours is 4%. (Sponsor)

The sites that block scrapers are exactly the ones data teams need, but most benchmarks ignore those. String's open-source benchmark looks at the hard cases.

The results: String's Web Access API hit 96% success on the hardest targets. Browserbase? 51%. Nimble? 59%. Oxylabs? 69%. Firecrawl? 72%.

Real Amazon prices, live Zillow comps, Indeed postings, if it's online, String will reach it, and we'll do it faster and cheaper than your current vendor.

Fetch, search, map, and browser endpoints. Use an API key or MCP.

Try String for free here

Want to brainstorm your hardest use cases? Book time with us and we'll solve them together

📱

Deep Dives

Accelerating Apache Spark Queries (and Iceberg Rust Development) with Apache DataFusion Comet (7 minute read)

DataFusion Comet speeds Spark reads on Iceberg by keeping Java planning while using Iceberg Rust for native Arrow execution. On a 3 TB TPC-DS benchmark, it accelerated 102 of 103 queries and cut total runtime by about 40%, while its use of Iceberg Java's Spark test suites has helped drive more than 40 Iceberg Rust pull requests.
From OTEL to SLMs: Distilling Frontier Model Behaviour from Production Telemetry (45 minute video)

OpenTelemetry traces can convert production AI-agent feedback into training labels for local 7B-13B coding models. Fine-tuned within hours, they can reach roughly 80-85% of frontier-model quality while lowering inference costs, preserving privacy, and scaling across the enterprise.
Exploring Hierarchical Interest Representation For Meta Ads Deep Funnel Optimization (8 minute read)

Meta compresses its graph of billions of users and advertiser entities into clusters of latent interests, enriches them with LLM-processed ad content, and embeds everything in one shared space where any user or entity can be scored against any other. Learning from broad graph structure instead of only rare conversion events gives the deep funnel a usable signal, and the embeddings feed Meta's ad ranking models like GEM and Andromeda.
🚀

Opinions & Advice

Data Modeling isn't Dead, You Just Stopped Doing It with Joe Reis (65 minute video)

Data modeling is still essential, especially in the AI era, because clear structures, standards, and context help teams avoid bad data, wasted work, and constant firefighting. Modern data professionals need a flexible “mixed model” approach that uses the right modeling style for the problem rather than forcing everything into one method.
Agents think in milliseconds, legacy infrastructure doesn't. LinkedIn, Walmart, and Zendesk shared how they closed the gap at VB Transform 2026 (4 minute read)

Enterprise agents are constrained more by legacy infrastructure than model quality. LinkedIn pre-provisioned containers and shifted 80% of orchestration to deterministic code, Walmart governed duplicate internal agents, and Zendesk strengthened data pipelines for its 20B conversations. The lesson is to invest in evals, own the agent harness, and keep workloads portable across models.
AI Agents Need Data Product Context Not More RAG (11 minute read)

Enterprise AI agents need governed context, not just vector search: retrieval can find relevant passages, but it does not preserve source boundaries, business meaning, ownership, or usage restrictions. The proposed approach compiles one governed data-product model into multiple representations: graph, Markdown wiki, portable OKF bundle, and compact sidecars such as TOON/GCF. AI readiness depends on connected, trusted context and repeatable operations rather than catalog presence alone.
💻

Launches & Tools

😥 CEO waiting for data? Cube agents deliver so your team can focus on thinking (Sponsor)

Asking a data question is still a ticket, a queue, and a week of bogged down data teams. Cube hands that workflow to AI agents who model, explore, and build reports - working from one governed set of definitions and single source of truth. Brex, Wix, and Patagonia run on Cube. Get started for free
Introducing Apache Spark 4.2 (6 minute read)

Apache Spark 4.2 adds metric views for defining business metrics, SQL vector similarity search, native geospatial types, and change data capture queries through a new CHANGES clause. Spark Connect and default Arrow optimized Python execution make the engine easier to reach. The release bundles over 1,900 commits from more than 260 contributors.
Ontology Playground (GitHub Repo)

Microsoft's Ontology Playground is a free, open-source tool for visually exploring, building, importing and sharing ontologies for Microsoft Fabric IQ.
Postgres 19 Compression: from pglz to LZ4 (8 minute read)

Postgres 19 plans to switch default TOAST compression from pglz to LZ4, while keeping the same unified compression framework across heap storage, TOAST, and B-tree indexes. Variable-length types like TEXT, VARCHAR, BYTEA, and JSONB are compressed automatically via varlena metadata, with rows targeting ~2 KB before spilling out-of-line to TOAST. LZ4 is much faster than pglz in tests, often compresses better, and still fails fast on incompressible data.
Experience Graphs: The Data Foundation for Self-Improving Agents (28 minute read)

Trellis stores agent artifacts, rewards, search data, and causal history as persistent experience graphs accessible through SQL, graph, vector, and temporal queries. Meta found that reusing past experience matched performance about 10x faster and reduced successful-solution token costs by 52%, though excessive memory reuse reduced exploration.
🎁

Miscellaneous

In-House LLM Serving at Netflix (7 minute read)

Netflix runs LLM serving end-to-end in-house, using a unified JVM-based platform with gRPC plus an OpenAI-compatible HTTP API, rather than relying on hosted APIs. The stack standardizes deployment through Triton-backed Model Scoring Service, using vLLM as the paved-path engine after re-benchmarking. It also adds custom packaging, zero-downtime rollouts, and patched guided decoding to prevent malformed JSON output.
From Weeks to a Day: How We Made LLM Evaluation Fast Enough to Iterate on (10 minute read)

Airbnb cut LLM evaluation cycles from weeks to under a day by caching generated references and judge scores, since over half of model outputs across candidates were identical strings and judge drift of about 1% per run. Small LoRA adapters with rank under 50 train in under an hour on a single GPU for same day fixes, and a final validation stage runs sampled traffic to catch failures.

Quick Links

Databricks hits $188B valuation, extending its run as AI's favorite second act (4 minute read)

Databricks is reportedly raising $3 billion at a $188 billion valuation after successfully repositioning itself from a data platform into a major enterprise AI provider.
Dremio's Exit Is the Clearest Sign Yet That Lakehouse-Only Won't Survive AI (3 minute read)

Dremio's sale to SAP is presented as evidence that data platforms built around a single lakehouse will struggle with AI's demands for broad reach, freshness and unpredictable querying.

Love TLDR? Tell your friends and get rewards!

Share your referral link below with friends to get free TLDR swag!
Track your referrals here.

Want to advertise in TLDR? 📰

If your company is interested in reaching an audience of data engineering professionals and decision makers, you may want to advertise with us.

Want to work at TLDR? 💼

Apply here, create your own role or send a friend's resume to jobs@tldr.tech and get $1k if we hire them! TLDR is one of Inc.'s Best Bootstrapped businesses of 2025.

If you have any comments or feedback, just respond to this email!

Thanks for reading,
Joel Van Veluwen, Tzu-Ruey Ching & Remi Turpaud


Manage your subscriptions to our other newsletters on tech, startups, and programming. Or if TLDR Data isn't for you, please unsubscribe.

Post a Comment

0 Comments