TLDR Data #100 - A year in review (Website) 100 issues. One year of TLDR Data. This year was shaped by faster analytical engines, pragmatic database architectures, and real-world tradeoffs, including those around open data formats. We dug into AI governance and infrastructure shaping today, and context engineering and semantic layer patterns emerging next. Dive into the full review for the highlights and deeper trends. Thanks for reading, building, and sharing with us. | Sampling: the philosopher's stone of distributed tracing (11 minute read) Distributed tracing, powered by OpenTelemetry, is essential for modern observability but generates massive volumes of span data that outpace current querying capabilities. Sampling (primarily via head or tail methods) remains vital to control scale and cost: head sampling offers deterministic, stateless selection, while tail sampling enables context-aware retention but introduces significant architectural and operational complexity, particularly in multi-zone environments and accurate RED metric production. Practical implementations require smart routing, metric extraction before sampling, and a clear understanding of trade-offs. Emerging solutions like disk-based buffering and exemplar-based sampling offer incremental improvements. | From firefighting to building: How AI agents restored our team's core productivity (18 minute read) Grab's Analytics Data Warehouse team deployed a multi-agent AI system to autonomously resolve up to 40% of repetitive user queries, reclaiming hundreds of engineering hours monthly across their 1,000+ users and 15,000+ tables. Specialized agents—handling tasks from code enhancement to data lineage investigation—replace manual triage, while layered safeguards and a human-in-the-loop process ensure security, data quality, and user trust. | Migrating Etsy's database sharding to Vitess (9 minute read) Etsy migrated from a legacy MySQL sharding setup (single unsharded index database as a central lookup across 1,000+ tables) to Vitess, eliminating the index database (a single point of failure), automating scaling (from months to days), and abstracting sharding complexity from developers. Its team implemented custom vindexes, rolled out incrementally table-by-table with quick rollbacks, and achieved zero-downtime cutover with no massive data movement. | | The Living Graph: Holons and the Four-Graph Model (21 minute read) The holonic model introduces a four-layer architecture for RDF knowledge graphs: interior, boundary/membrane (via SHACL), projection, and context. This enables true encapsulation, boundary governance, and contextual relationships. Each holon is both a self-contained entity and part of a holarchy, facilitating multi-resolution querying, provenance tracking, and federated authority. Portals enable controlled, annotated traversal between holons, solving challenges in representing containment, authority, and navigation that flat RDF graphs cannot address. This principled structure brings organization, scalability, and semantic richness crucial for complex, multi-domain graph systems. | The Question Is the Contract (17 minute read) Competency questions are explicit, domain-specific queries that a system must answer. Systems built without first defining and maintaining these questions often yield inaccurate or incomplete results, because architectural decisions default to implicit assumptions rather than clear requirements. Leveraging competency questions provides a rigorous, testable framework that guides schema design, validates coverage, and ensures traceable, actionable answers in everything from vector databases and knowledge graphs to retrieval-augmented generation pipelines. | AI is redrawing the database market (13 minute read) AI workloads are forcing a convergence of previously separate domains (real-time analytics, data warehousing, and observability) into a single high-concurrency, low-latency, full-fidelity data platform, as legacy batch-oriented systems can't handle bursts of concurrent interactive queries, massive unsampled data needs, or real-time freshness without failing. | | Introducing the Apache Airflow Registry (2 minute read) The Apache Airflow Registry offers a centralized, searchable catalog of 98 providers and over 1,600 modules, including 848 operators, 298 hooks, and 372 modules for Amazon alone. Features include instant search (Cmd+K), provider-specific pages with one-click install commands, integrated connection builders generating URI, JSON, or Env Var formats, and a JSON API for programmatic integration with tools and IDEs. This streamlines module discovery, configuration, and automation. | How ROOST is Advancing Online Safety (7 minute read) Discord donated its internal high-performance rules engine, Osprey (now production-ready and community-improved), to ROOST. It processes real-time platform events (logins, messages, and account changes) to detect threats instantly, evaluate thousands of rules across actions, flag suspicious activity, and support investigations with managed services from partners like Musubi and Zentropi. | QCon London 2026: Introducing Tansu.io - Rethinking Kafka for Lean Operations (4 minute read) Tansu.io introduces a Kafka-compatible, stateless messaging broker where durability shifts entirely to external storage, radically reducing broker memory footprint (~20MB) and enabling instant scale-to-zero deployments within 10ms. With pluggable storage backends (S3, SQLite, or direct Postgres), Tansu simplifies streaming pipelines by writing validated records (Avro, JSON, or Protobuf) directly to open table formats like Iceberg and Delta Lake, and eliminates the transactional outbox. Code available on GitHub. | | Neoclouds and Why AI Compute Demand has Outgrown Hyperscaler Balance Sheets (7 minute read) Hyperscalers' growth (now near cash flow ceilings) cannot keep pace, making neoclouds indispensable overflow valves in a structurally supply-constrained, geography-bound market. AI compute demand has caused Azure's backlog to surge 1,150% to $625B, yet even Microsoft cannot self-supply, driving rapid partnerships with neoclouds like CoreWeave and Crusoe, whose combined GPU commitments now exceed $131B. Large AI labs such as OpenAI and Anthropic now diversify away from single-cloud dependence, while neoclouds provide 10–20% of AI capex. | 5 Production Scaling Challenges for Agentic AI in 2026 (4 minute read) Moving autonomous agentic AI systems from impressive demos to reliable production remains far harder than traditional ML, due to amplified issues in coordination, visibility, economics, testing, and control. Probabilistic outputs break deterministic testing. Emerging LLM-as-judge and simulations help, but still need heavy human oversight and lack standardization. | | | Love TLDR? Tell your friends and get rewards! | | Share your referral link below with friends to get free TLDR swag! | | | | Track your referrals here. | | | |
0 Comments