rustc_codegen_jvm: Rust compiler backend to emit JVM bytecode
Compilers
Why Janet?
Artificial Intelligence
How small businesses can leverage AI
This article is from Making AI Work, MIT Technology Review’s limited-run newsletter examining how to apply LLMs across industries. To receive it in your inbox,sign up here. From accounting to design to market research and product development, there’s a staggering breadth of skills needed to run a business. A large company can hire experts to…
AI Agents
Fragments: June 2
Greg Wilson has noticed that lots of folks are using dodgy metrics to figure out if AI tools are worth their costs. Would you measure lines of code generated, or tickets closed? Or would you send out a survey asking whether developers feel more productive? Each of those approaches is flawed in a different way; He lists lots of common metrics, and why they are flawed. Sadly he doesn’t give any suggestions on what would be better. In my view, since we cannot measure productivity, any me...
Not Every Byte Gets a Vote
Muxcard, a dyi credit card size computer
Web Application Security
ChatGPhish: The Page Is the Payload
CQL: Categorical Databases
Every byte matters
AI Agents
How to Build a Shitty Robot
Show HN: AI Simulaionen Based on FEP
Meta
I Love Meta Platforms
AI Coding Tools
strace-ui, Bonsai_term, and the TUI renaissance
AI Inference
Mellum2 Goes Open Source: A Fast Model for AI Workflows
Large Language Models
AnyEdit++: Adaptive Long-Form Knowledge Editing via Bayesian Surprise
arXiv:2606.01053v1 Announce Type: new Abstract: Editing complex, long-form knowledge in Large Language Models remains a significant challenge due to the difficulty of maintaining generation coherence. Existing autoregressive methods like AnyEdit alleviate length constraints but rely on Fixed-window Chunking, which disregards logical structure and compromises consistency. To address this, we present AnyEdit++, a structure-aware framework incorporating Bayes-Chunk, an adaptive segmentation mec...
Generative AI
TextFake: Benchmarking AI-Generated Image Detection on Text-Rich Images
arXiv:2606.01050v1 Announce Type: new Abstract: Recent AI-generated image (AIGI) detectors perform well on natural-image benchmarks, but their behavior on text-rich forgeries, such as fabricated screenshots, documents, and news pages prevalent in misinformation, remains untested. We introduce TextFake, a 20,000-image benchmark for text-rich AIGI detection spanning 28 languages, 4 topic categories, and 2 scene modalities. Fake images are synthesized via a four-stage pipeline that annotates re...
Embeddings
PMC-InterCPT: Rethinking Biomedical Interleaved Data for Multimodal Continued Pretraining
arXiv:2606.01049v1 Announce Type: new Abstract: Large-scale biomedical image-text datasets extracted from scientific literature provide valuable resources for medical multimodal model training. These datasets are commonly organized as image-caption pairs; however, figure captions are often short, context-dependent, and only partially informative without the surrounding article text. At the same time, large-scale automatic extraction introduces structural noise such as missing captions, resid...
Diffusion Models
Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation
arXiv:2606.01048v1 Announce Type: new Abstract: We propose Decoupled Residual Denoising Diffusion models (DRDD) for unified and data-efficient image-to-image (I2I) translation. While diffusion models have advanced I2I translation in terms of quality and diversity, we uncover a previously under-explored property in diffusion models. Crucially, beyond its conventional role of manifold lifting (i.e., moving data off low-dimensional manifolds), injecting Gaussian noise facilitates domain harmoni...
Robotics
Learning Multi-Modal Trajectory Policies for Data-Efficient Robotic Manipulation
arXiv:2606.01047v1 Announce Type: new Abstract: Robotic manipulation requires the effective integration of heterogeneous inputs, including visual observations, language instructions, and trajectory representations, to generate accurate actions. Existing transformer-based policies typically process these heterogeneous modalities within a shared parameter space, which often leads to modality interference and inefficient representation learning, especially in data-scarce scenarios. While Mixtur...
LLM Evaluation
TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents
arXiv:2606.01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existing benchmarks' limitations: 1) overemphasis on constraint compliance, neglecting multi-dimensional qualities like spatio-temporal cost; 2) datasets lacking real-world authenticity and coverage in key areas (e.g., lodging, transport); and 3) isolated daily plan assessments that miss critical deta...
Large Language Models
Child-directed speech facilitates production, not comprehension, in BabyLMs
arXiv:2606.01045v1 Announce Type: new Abstract: Recent studies suggest that child-directed speech is not conducive to language learning in BabyLMs. However, current evaluations focus predominantly on comprehension and not production, which is central to usage-based theories of language acquisition which argue how CDS facilitates early language use through constructional ''frames'' (frequent lexical patterns with open slots). We introduce a novel generation-based evaluation inspired by such t...
Diffusion Models
Ask4VG: Risk-Aware Question Selection for Reducing Prior-Driven Answers in Medical VQA
arXiv:2606.01044v1 Announce Type: new Abstract: Medical visual question answering requires models to ground their responses in image evidence, because visually unsupported answers can mislead downstream interpretation. However, many medical VQA questions are generic, template-like, or highly similar in form, which can encourage models to learn question-answer shortcuts instead of image-dependent reasoning and thereby increase the risk of hallucinated responses. We propose Ask4VG, a label-fre...
Computational Biology
Plausibility Is Not Prediction: Contrastive Evidence for LLM-Based Cellular Perturbation Reasoning
arXiv:2606.01042v1 Announce Type: new Abstract: Perturbation experiments are central to understanding cellular mechanisms, but remain costly and sparse, motivating prediction of gene expression responses for unobserved conditions. A promising recent direction leverages large language models (LLMs) as "virtual cell" simulators-using stepwise, knowledge-grounded mechanistic reasoning to infer differential expression-pointing toward an interpretable, knowledge-driven paradigm that transcends pu...
EasyOS built with Xlibre
Stripe
I Got $4.84 from a Class Action and They Didn't Want Me to Have It
Parallel Reconstruction of Lawful TLS Wiretapping
Microsoft AI
Angry devs vow to flee GitHub Copilot as metered billing takes hold
Network Security
U.S. Midterms Have a Cyber Problem, but It's Not at the Ballot Box
AI Psychosis
Remember This
NVIDIA
How the hell is Groq raising more money?
Stripe
macOS Needs Its Grid Back
Fintech
Squillions: How Money Laundering Won
Crystal Nights by Greg Egan
Constant Q Transform – A Visual Guide
A Structure-Aware Fuzzing Experiment
Chipotlai Max
Query Planning
Christophe Pettus: All Your GUCs in a Row: cpu_index_tuple_cost, cpu_operator_cost, and cpu_tuple_cost
cpu_tuple_cost, cpu_index_tuple_cost, and cpu_operator_cost are three of the constants the planner uses to price a query, and the single most useful thing to know about all three is that you should almost certainly never change them. The rest of this post is why. PostgreSQL’s planner does not est…
Software Engineering
What's gonna happen to software engineers?
Startups and Venture
Can the stockmarket swallow Anthropic, SpaceX and OpenAI?
Privacy
Age verification for social media, the beginning of the end for a free internet?
A new way to build chips: Sequentially stacking silicon to extend Moore's Law
The Frame Problem
Lid-lifting Kiwi author forced to sit in silence at writers' festival
Networking Protocols
Medium Access Control Protocols
Show HN: DepsGuard – one command to harden NPM/pnpm/yarn/bun/uv configs
OpenAI
OpenAI frontier models and Codex are now available on AWS
Startups and Venture
How to make the Startup Battlefield Top 20 — and what every company gets regardless
Google
Alphabet plans to raise $80 billion to pay for AI buildout
Building a custom mount for a telescoping webcam
Teaching AI to Simulate Nature, Faster
AI frameworks called neural operators have been applied in various scientific and engineering domains.
Software Engineering
Install web apps with the new HTML install element
AWS
Scaling oncology patient support: How New York Cancer and Blood Specialists transformed customer experience with AWS and Pronetx, now part of Caylent
This post details how NYCBS partnered with Amazon Web Services (AWS) and AWS partner Pronetx (now part of Caylent) to migrate to Amazon Connect Customer, the AWS cloud contact center service. The migration delivered a 54 percent improvement in patient enrollment and transformed the way NYCBS connects with the patients who need them most.
Defense Tech
Defense tech darling Mach Industries hits $1.8B valuation, a 4x jump in a year
AI Inference
Get started with OpenAI GPT-5.5, GPT-5.4 models, and Codex on Amazon Bedrock
OpenAI frontier models GPT-5.5 and GPT-5.4, and Codex, the OpenAI coding agent, are now generally available on Amazon Bedrock. Deploy frontier models on Bedrock's high performance inference engine with built-in security, governance, and pay-per-token pricing.
NVIDIA
Nvidia chases $200B CPU market with AI agent PCs from Microsoft, Dell, and HP
I missed Network integrated tools on Windows so I built a Linux equivalent
Startups and Venture
From the stage to the future: Where are Startup Battlefield’s alumni now?
Show HN: Textile – A desktop app for weaving together bits of text
NVIDIA
Michael Burry Just Called Nvidia's SpaceX Chip Deal 'Fugazi.'
Florida sues OpenAI and Sam Altman over AI risks
AI Psychosis
Florida sues OpenAI, Sam Altman, in first-of-its-kind lawsuit over violent incidents
MySQL
Migrating Etsy’s database sharding to Vitess
Etsy has maintained a sharded MySQL architecture since around 2010. This database cluster contains most of Etsy’s online data and is made up of ~1,000 tables distributed across ~1,000 shards. Over the last 16 years, it has grown significantly: combined, these tables have over 425 TB of data and receive roughly 1.7 million requests per second. Etsy engineers access our MySQL data through a proprietary object-relational mapping (ORM). The ORM has a corresponding model for each MySQL table. W...
Artificial Intelligence
Making Ads Count: Using MMoE and Auxiliary Tasks to Better Connect Buyers & Sellers
When buyers search on Etsy, they need to quickly and easily find the perfect item. At the same time, sellers need to be confident their unique products are being seen by the right customers. Our Ads Search ranking model, which is built on a multitask learning foundation, is the critical link in this connection. Recently, we identified an opportunity to drive more meaningful buyer engagement by enhancing our model’s ability to predict purchase intent. We achieved this via a dual-pronged impr...
RAG
Shaping Product Understanding with Contrastive Reinforcement Learning
Etsy’s marketplace is defined by the creativity and craftsmanship of our sellers and the hundreds of millions of highly diverse products they offer. You can find silversmiths who cold-forge recycled sterling silver, weavers who dye raw fleece with indigo and black walnut, and ceramicists who throw stoneware on a kick wheel. These details define each product and often determine whether it matches a buyer’s taste, style, and interests. Sometimes buyers know exactly what they want, searching...
Build a Basic AI Agent from Scratch: Tools
Florida AG files lawsuit against OpenAI, CEO Sam Altman for deceptive practices
Homelab & Self-Hosting
How we reduced core unit boot time from hours to minutes
We investigated why firmware updates were causing our core servers to take four hours to reboot. By diving into UEFI data structures and iPXE automation, we eliminated unnecessary timeouts and cut boot times back down to minutes.
Hackers hijacked Instagram accounts by tricking Meta AI support chatbot into granting access
Qwen3.7-Plus: Multimodal Agent Intelligence
iSCSI CHAP: Heap Buffer Overflow in the Linux Kernel
Algorithms and Data Structures
Stealing from Biologists to Compile Haskell Faster
Better Than the Truth: From AI Hallucinations to Imaginations
I see enormous potential in what I call AI's “deliberate imaginations” of reality, or possibilities beyond reality.
EVs and Transportation
Water access is now a risk factor in SpaceX’s IPO
Checking Assembly with Z3
X / Twitter
AI Grifters Are Making Anti-Data Center Slop With AI
There are hundreds of anti-data center Facebook pages churning out AI-generated slopaganda.
Privacy
We Sued ICE to Get Its Spyware Contract. The Agency Is Redacting Essentially Everything
Paragon's software is capable of remotely breaking into phones and accessing messages from encrypted messaging apps. Our lawsuit aims to pry records about it from ICE.
Amazon Shuts Down Internal AI Leaderboard After Employees Cheated
Employees admitted to 404 Media they had cheated to climb the leaderboard's ranks.
Superintelligence: The Idea That Eats Smart People (2016)
Computational Biology