2-D Mathematical Curves
OpenAI can't build working RSS feeds
MySQL
Migrating Etsy’s database sharding to Vitess
Etsy has maintained a sharded MySQL architecture since around 2010. This database cluster contains most of Etsy’s online data and is made up of ~1,000 tables distributed across ~1,000 shards. Over the last 16 years, it has grown significantly: combined, these tables have over 425 TB of data and receive roughly 1.7 million requests per second. Etsy engineers access our MySQL data through a proprietary object-relational mapping (ORM). The ORM has a corresponding model for each MySQL table. W...
Artificial Intelligence
Making Ads Count: Using MMoE and Auxiliary Tasks to Better Connect Buyers & Sellers
When buyers search on Etsy, they need to quickly and easily find the perfect item. At the same time, sellers need to be confident their unique products are being seen by the right customers. Our Ads Search ranking model, which is built on a multitask learning foundation, is the critical link in this connection. Recently, we identified an opportunity to drive more meaningful buyer engagement by enhancing our model’s ability to predict purchase intent. We achieved this via a dual-pronged impr...
Artificial Intelligence
The Pulse of India’s Supply Chain: How Flipkart Plans for 500 Million Customers — Part 1
In the world of e-commerce, a click of a “Buy Now” button starts an invisible race. Across India (metro cities or remote villages) millions of customers expect their packages to arrive with a speed that feels like magic.But behind that “magic” is a monumental logistical puzzle. At Flipkart, we serve over 500 million registered users. We manage 150 million products. On an average day, we handle 4 million shipments. During our biggest sales events, like the Big Billion Days (BBD), that...
Train Your Own LLM from Scratch
Search Engines
What (un)exactly do you mean by semantic search?
Ryan welcomes Bryan O’Grady, Head of Field Research and Solutions Architecture at Qdrant, to discuss the differences between traditional text search engines powered by Lucene and modern vector databases, when vector search’s exact-match needs work for things like logs and security analytics and when semantic search works for user-facing discovery and non-exact results, and how Qdrant is growing into video embeddings and local-agent contexts.
AI Coding Tools
An LLM agent that runs on any Linux box
Machine Learning
Sparse Regression under Correlation and Weak Signals: A Reproducible Benchmark of Classical and Bayesian Methods
arXiv:2605.00835v1 Announce Type: new Abstract: Choosing between classical and Bayesian sparse regression methods involves a real trade-off: penalized estimators like Lasso run in milliseconds but give no uncertainty estimates,while Horseshoe and Spike-and-Slab priors produce full posteriors but need MCMC chains that take minutes per fit.Surprisingly few studies compare these two families head-to-head under the conditions that actually make sparse regression hard -- correlated features, weak...
Linear Algebra
Polynomial-Time Optimal Group Selection via the Double-Commutator Eigenvalue Problem
arXiv:2605.00834v1 Announce Type: new Abstract: The algebraic diversity framework replaces temporal averaging over multiple observations with algebraic group action on a single observation for second-order statistical estimation. The central open problem in this framework is $\textit{group selection}$: given an $M$-dimensional observation with unknown covariance structure, find the finite group whose spectral decomposition best matches the covariance. Naive enumeration of all subgroups of th...
Artificial Intelligence
Agentopic: A Generative AI Agent Workflow for Explainable Topic Modeling
arXiv:2605.00833v1 Announce Type: new Abstract: Agentopic is a novel agent-based workflow for explainable topic modeling that leverages the reasoning capabilities of Large Language Models (LLMs). Existing topic modeling approaches such as Latent Dirichlet Allocation (LDA) and BERTopic often lack transparency on how topics are assigned or grouped. Agentopic addresses this by using multiple agents that collaboratively perform topic identification, validation, hierarchical grouping, and natural...
Generative AI
Synthetic Designed Experiments for Diagnosing Vision Model Failure
arXiv:2605.00832v1 Announce Type: new Abstract: Current synthetic data pipelines for computer vision generate images without diagnosing what the downstream model actually needs. This open-loop paradigm treats synthetic data as cheap real data, randomly sampling the generator's output space and hoping to cover the model's failure modes. We argue this fundamentally misuses synthetic data's unique property: the controllable, independent variation of scene factors.Drawing on the statistical theo...
AI Inference
GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving
arXiv:2605.00831v1 Announce Type: new Abstract: The rise of million-token, agent-based applications has placed unprecedented demands on large language model (LLM) inference services. The long-running nature of these tasks increases their susceptibility to hardware and software faults, leading to costly job failures, wasted resources, and degraded user experience. The stateful key-value (KV) cache, which grows with the sequence length, presents a central challenge as it is a critical and vuln...
GPU and Parallel Computing
Efficient Accelerated Graph Edit Distance Computation on GPU
arXiv:2605.00830v1 Announce Type: new Abstract: Graph representation is a powerful abstraction of real-world objects and relations. Computing the Graph Edit Distance (GED) between graphs is critical in domains such as bioinformatics, machine learning, and pattern recognition. GED measures the minimum number of edit operations required to transform one graph into another. However, the high computational complexity of optimal and near-optimal methods limits their applicability to large-scale g...
Monitoring and Alerting
LLM-based uncertainty assessment of social media situational signals for crisis reporting
arXiv:2605.00829v1 Announce Type: new Abstract: Social media has become a critical source of situational awareness during disasters, providing real-time insights into evolving impacts and emerging needs. To support crisis response at scale, recent work has increasingly leveraged large language models (LLMs) to automatically classify and summarize situational information from social media streams. However, existing approaches implicitly assume that extracted situational claims are equally pla...
Blockchain
Canonical LST: A Protocol-Native Liquid Staking Solution for Tezos
arXiv:2605.00828v1 Announce Type: new Abstract: Canonical LST (sTEZ) is an enshrined, protocol-native mechanism designed to mitigate the centralization risks associated with liquid staking intermediaries. Intended to complement direct staking rather than replace it, Canonical LST provides a neutral, public alternative managed directly by the Tezos protocol. It allows any tez holder to participate in aggregated staking without reliance on third-party operators. sTEZ follows an accrual-based d...
AI Agents
Separating Intelligence from Execution: A Workflow Engine for the Model Context Protocol
arXiv:2605.00827v1 Announce Type: new Abstract: Large Language Model (LLM) agents increasingly interact with external systems through tool-calling protocols such as the Model Context Protocol (MCP). In prevailing architectures, the agent must reason about every tool invocation in every session, consuming tokens proportional to the number of actions performed--even when the task has been solved before. We present the MCP Workflow Engine, a novel MCP-native orchestration layer that decouples i...
RAG
Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis
arXiv:2605.00826v1 Announce Type: new Abstract: Text-to-video retrieval enables users to find relevant video content using natural language queries, a task that has grown increasingly important with the rapid expansion of online video. Over the past six years, research has produced numerous methods, such as dual encoders, attention-driven models, and multimodal fusion approaches; however, fundamental questions remain about model behavior, dataset influence, and query difficulty. In this work...
Linux, Windows or macOS: Which Operating System to Use in 2026?
CVE-2026-31431: Copy Fail vs. rootless containers
Apple
Apple Explores Using Intel and Samsung to Build Main Device Chips in the US
Shelley: Mobile-friendly, web-based, multi-modal, single-user coding agent
EVs and Transportation
The Car That Watches You Back: The Advertising Infrastructure of Modern Cars
NVIDIA
As workers worry about AI, Nvidia’s Jensen Huang says AI is ‘creating an enormous number of jobs’
PostgreSQL
PgQue: Two Snapshots and a Diff
Privacy
Meta, TikTok Recv Personal Data from Health Exchanges Alarming Privacy Experts
Database Administration and Tooling
pgxbackup: Continuity Support for pgBackRest
Artificial Intelligence
What I'm Hearing About Cognitive Debt (So Far)
When Networking Doesn't Work
SprintiQ – open-source sprint planning for Claude Code
Bun is being ported from Zig to Rust
Embedded Rust or C Firmware? Lessons from an Industrial Microcontroller Use Case with Ariel OS
xAI
Y Combinator's Stake in OpenAI (0.6%)
Artificial Intelligence
Researchers Asked LLMs for Strategic Advice. They Got "Trendslop" in Return
WebAssembly
The state of ARM64 on Windows in 2026
Amazon
Your Dinner Got Worse On Purpose
Show HN: I Built a Museum Exhibit
Browser Internals
Suspected YouTube bug spikes RAM over 7gbs users report lag and frozen tabs
Artificial Intelligence
What do we lose when AI does our work?
Testing
Agent Skills
Apple confirms iOS 26.5 Messages app adds RCS end-to-end encryption
Podman rootless containers and the Copy Fail exploit
CVE and Exploit Development
U.S. government warns of severe CopyFail bug affecting major versions of Linux
Apple
Testing macOS on the Apple Network Server 2.0 ROMs
xAI
OpenAI’s cozy partner Cerebras is on track for a blockbuster IPO
KDE Union: Spring 2026 Update
Computational Complexity
Transformers Are Inherently Succinct
Claude Is Dead
Infrastructure as Code
Welcome to Gas City
Frizbee is a tool you may throw a tag at and it comes back with a checksum
Stripe
Formatting a 25M-line codebase overnight
Security Advisory: Local privilege escalation in Lix and Nix
How AI is Changing Programming Language Usage
The impact of AI on software development involves more than code creation tools.
Artificial Intelligence in Physics: ‘The Influence is Huge’
Research groups have been using AI to address outstanding problems that have remained out of reach by traditional computational methods.
Should There Be Limitations on Post-Mortem Avatars?
Ghostbots may be comforting to some but others say they could disrupt the grieving process.
Why Periodic Testing Fails Modern Apps
Why periodic security testing fails modern apps, and what a stronger approach looks like in practice.
Real-Time Data Processing in UAV Systems
Drones can assist researchers investigating interactions between computational and real-life dynamics, blending theory and practical computer science in physically interactive contexts.
Yourdon Had a Point
Considering the notion that "international competition will put American programmers out of work.”
Empowerment over Automation in AI Summarization of Online Reviews
AI-generated review summaries must enable awareness, comparison, and control so that users remain active interpreters rather than silent recipients of algorithmic consensus.
Connecting Software and Manufacturing Through APIs
APIs connect software and manufacturing systems so that they can operate as one.
The Hidden Privacy Cost of Screenshot Collection in Computing Research
It is time for lawmakers and institutional review boards to reckon with screenshot collection’s privacy problem.
Connection Pooling
PGKeeper: Building the Bouncer We Needed for Postgres | Figma Blog
WebRTC
How OpenAI delivers low-latency voice AI at scale
CVE and Exploit Development
Why a Decade of Writing Detection Logic Makes the Mythos Exploit Numbers Less Scary
CVE and Exploit Development
uutils coreutils CVEs
Google is discontinuing its free web search index for developers
Anthropic
White House Considers Vetting A.I. Models Before They Are Released
Google
DHS Demanded Google Data on Canadian's Activity, Location over Anti-ICE Posts
Show HN: nfsdiag - a NFS diagnostic application
Release v0.9.0 · Foxboron/ssh-tpm-agent
Artificial Intelligence
Image AI models now drive app growth, beating chatbot upgrades
Links to CSS colour palettes
A while back I decided to stop using Tailwind for new projects and to just write vanilla CSS instead. But one thing I missed about Tailwind was the colour palette (here as CSS). If I wanted a light blue I could just use blue-100 and if I didn’t like it maybe try blue-200 or blue-50. I’m not very good with colours so it makes a big difference to me to have a reasonable colour palette that somebody who is better at colour than me has thought about. But I’m also a little tired of those Tai...
Oasis Linux
Securing a DoD Contractor: Finding a Multi-Tenant Authorization Vulnerability
Anthropic
Usage-based pricing killing your vibe, here's how to roll your own local AI
Software Engineering
Let's Talk about LLMs
Mobile Development
Discord Patch Notes: May 4, 2026
Check out the finer details of the more technical fixes implemented into Discord recently.
Web Application Security
Hackers are still exploiting the cPanel bug to gain control of thousands of websites
Space Tech