Zotero Catalogue

A public bibliography of papers, books, and resources.

793 items Sort: Newest Oldest Title A–Z Title Z–A

YI78CHSN
blogPost
The MIT Press Reader
2026
Saved 2026-07-30
Decades before recommendation algorithms, video stores served as social hubs where conversation and camaraderie became a form of taste-making.
VBRQF6UV
videoRecording
Holley Muraco at Something Wild
2026
Saved 2026-07-30
We are currently fundraising for an animal care building that will become the core of Something Wild's rescue mission and will allow us to take on more animals and volunteers. Please consider donating on paypal, venmo, by cheque or here on our YouTube fundraiser. Every penny counts! You can also support us by buying adorable beaver merch from our online store.
GJD3RHGX
blogPost
Nathan Lalonde
2026
Saved 2026-07-30
Authors: Ahmed ElKady, Ahmed Y. Radwan, Shaina Raza Mixture-of-Experts (MoE) models carry far more parameters than they activate. For each token that passes through, a learned router selects only a […]
NDFBMRTA
webpage
Alexey Guzey
Saved 2026-07-30
Note: also check out my weekly “Best of Twitter” newsletter How to start using Twitter Sign up Check out the 20 people I like the most tweets from and follow the ones you like the most. Read Lama Al Rajih’s Twitter guide Rules for making twitter a wholesome experience Mute people you don’t want to see on your feed. It’s that simple. It will take a week or two when you start but now my feed, for example, has ~0 politics and engagement bait. Cool people …
SPM5FCWV
webpage
Alexey Guzey
Saved 2026-07-30
Summary: in this post I explain why you should start a blog (to help others and to help yourself), what to write about, and how to start it. I hope to persuade you that you should start a blog even if you feel that you have nothing to say and even if almost nobody will read it. What to write about I looked over all of my writing and determined that it all originated from one of the following: I repeatedly gave the same advice to my friends Why You Should Join Twitter Right Now Every …
5SR5IXB6
webpage
Saved 2026-07-30
I turned 30 last week and a friend asked me if I'd figured out any life advice in the past decade worth passing on.  I'm somewhat hesitant to publish this because I think these lists usually seem...
MX7IML9E
videoRecording
Lee Walters
2020
Saved 2026-07-30
Cute lamb needs attention, Trying to hold the camera and pet the cute lamb at the same time 😂
AL3KCHSE
forumPost
Stanley Tang [@stanleytang]
2026
Saved 2026-07-30
D2NU36AK
webpage
Saved 2026-07-30
Parents want accountability after study went disastrously wrong, Science and Retraction Watch investigation reveals
EB9IEEET
forumPost
Ash Tilawat [@ashtilawat]
2026
Saved 2026-07-30
EINIX77Q
webpage
Saved 2026-07-30
Thread by @patio11: "Some people really benefit from hearing advice that everyone knows, for the same reason we keep schools open despite evein them having been taught before. In that spirit, here's some quick Things Many People Find Too Obvious To H […]"
RXFL5RYZ
webpage
Saved 2026-07-28
KUUT8ZWJ
journalArticle
JOE: a mobile, inverted pendulum
Felix Grasser, Aldo D'Arrigo, Silvio Colombi, Alfred Rufer
2002 · IEEE Transactions on Industrial Electronics
Saved 2026-07-28
The Industrial Electronics Laboratory at the Swiss Federal Institute of Technology, Lausanne, Switzerland, has built a prototype of a revolutionary two-wheeled vehicle. Due to its configuration with two coaxial wheels, each of which is coupled to a DC motor, the vehicle is able to do stationary U-turns. A control system, made up of two decoupled state-space controllers, pilots the motors so as to keep the system in equilibrium
3E7Q2AQP
webpage
Garry Tan
2026
Saved 2026-07-28
1.5% of SF’s population lives in city-funded housing. 26% of its drug OD deaths happen there.
UZTQPBFY
webpage
Alexi Gladstone
Saved 2026-07-28
Autoregressive and diffusion models are trained on a single step but inferenced on thousands. Why does this work at all?
4CRFPYF7
forumPost
EO [@eostudi0]
2026
Saved 2026-07-28
7VTHGG4Z
webpage
Bracket Bot Capstone
Saved 2026-07-27
Bill of Materials.
6CEMC64G
webpage
Saved 2026-07-27
R8DDG9TE
webpage
Saved 2026-07-27
NZDDRDS2
book
An introduction to the three volumes of Karl Marx's Capital
Michael Heinrich
2012 · Monthly review press
Saved 2026-07-27
W5SHKFIF
book
Karl Marx
Saved 2026-07-27
EY3JKYHX
journalArticle
The Question Concerning Technology
Martin Heidegger
Saved 2026-07-27
Z3PF8S87
book
The Lonely Man of Faith
Rabbi Joseph B. Soloveitchik, Joseph Dov Soloveitchik
2011 · Maggid Books
Saved 2026-07-27
The Lonely Man of Faith is a timeless philosophical essay by one of the twentieth century¿s greatest Jewish philosophers, Talmudic scholars, and religious leaders, Rabbi Joseph B. Soloveitchik. In this classic work, Rabbi Soloveitchik probes the inner experience of those who seek both redemptive closeness with God and creative engagement with the world. With characteristic brilliance and eloquence, he delineates the struggle of people of faith to navigate between seemingly contradictory aspects of the human condition: the spiritual and the material, the religious and the scientific, the covenantal and the majestic.
3TWDUTM5
webpage
Friedrich Hayek
Saved 2026-07-27
UXLTYU23
book
The Great Stagnation: How America Ate All the Low-hanging Fruit of Modern History, Got Sick, and Will (eventually) Feel Better
Tyler Cowen
2011 · Dutton
Saved 2026-07-27
Tyler Cowen's The Great Stagnation, the eSpecial heard round the world that ignited a firestorm of debate and redefined the nature of our economic malaise, is now-at last-a book. America has been through the biggest financial crisis since the great Depression, unemployment numbers are frightening, media wages have been flat since the 1970s, and it is common to expect that things will get worse before they get better. Certainly, the multidecade stagnation is not yet over. How will we get out of this mess? One political party tries to increase government spending even when we have no good plan for paying for ballooning programs like Medicare and Social Security. The other party seems to think tax cuts will raise revenue and has a record of creating bigger fiscal disasters that the first. Where does this madness come from? As Cowen argues, our economy has enjoyed low-hanging fruit since the seventeenth century: free land, immigrant labor, and powerful new technologies. But during the last forty years, the low-hanging fruit started disappearing, and we started pretending it was still there. We have failed to recognize that we are at a technological plateau. The fruit trees are barer than we want to believe. That's it. That is what has gone wrong and that is why our politics is crazy. Cowen reveals the underlying causes of our past prosperity and how we will generate it again. This is a passionate call for a new respect of scientific innovations that benefit not only the powerful elites, but humanity as a whole.
728M469B
book
Stubborn Attachments: A Vision for a Society of Free, Prosperous, and Responsible Individuals
Tyler Cowen
2018 · Stripe Press
Saved 2026-07-27
From a bestselling author and economist, a contemporary moral case for economic growth—and a dose of inspiration and optimism about our future possibilities.Growth is good. Throughout history, economic growth in particular has alleviated human misery, improved human happiness and opportunity, and lengthened human lives. Wealthier societies are more stable, offer better living standards, produce better medicines, and ensure greater autonomy, greater fulfillment, and more sources of fun. If we want to continue our trend of growth—and the overwhelmingly positive outcomes for societies that come with it—every individual must become more concerned with the welfare of those around us. So how do we proceed? Tyler Cowen, in a culmination of 20 years of thinking and research, provides a roadmap for moving forward. In Stubborn Attachments: A Vision for a Society of Free, Prosperous, and Responsible Individuals, he argues that our reason and common sense can help free us of the faulty ideas that hold us back as people and as a society, allowing us to set our sights on the long-term struggles that maximize sustainable economic growth while respecting human rights. Stubborn Attachments, at its heart, makes the contemporary moral case for economic growth, and delivers a great dose of inspiration and optimism about our future possibilities.
HHKXIX2F
journalArticle
Rationality: From AI to Zombies
Eliezer Yudkowsky
Saved 2026-07-27
5VHQNG5Q
book
Rationality: From AI to Zombies
Eliezer Yudkowsky
2015 · Machine Intelligence Research Institute
Saved 2026-07-27
USN7PF3I
book
Introduction to algorithms
Thomas H. Cormen
2009 · MIT Press
Saved 2026-07-27
FZDAZCKX
book
For a new liberty: the libertarian manifesto
Murray N. Rothbard, Llewellyn H., Jr Rockwell
2006 · Ludwig von Mises Inst.
Saved 2026-07-27
4Q9HTXBA
book
For a new liberty: the libertarian manifesto
Murray N. Rothbard, Llewellyn H., Jr Rockwell
2006 · Ludwig von Mises Inst.
Saved 2026-07-27
ZIMV72IY
book
For a new liberty: the libertarian manifesto
Murray N. Rothbard, Llewellyn H., Jr Rockwell
2006 · Ludwig von Mises Inst.
Saved 2026-07-27
22NAHZ4D
book
The myth of the rational voter: why democracies choose bad policies
Bryan Douglas Caplan
2008 · Princeton University Press
Saved 2026-07-27
"Caplan argues that voters continually elect politicians who either share their biases or else pretend to, resulting in bad policies winning again and again by popular demand. Calling into question our most basic assumptions about American politics, Caplan contends that democracy fails precisely because it does what voters want. Through an analysis of American's voting behavior and opinions on a range of economic issues, he makes the case that noneconomists suffer from four prevailing biases: they underestimate the wisdom of the market mechanism, distrust foreigners, undervalue the benefits of conserving labor, and pessimistically believe the economy is going from bad to worse. Caplan lays out several ways to make democratic government work better
PL3VX7YH
book
The socialist manifesto the case for radical politics in an era of extreme inequality
Bhaskar Sunkara
2020 · Verso
Saved 2026-07-27
8QKQ88IR
book
The long depression: how it happened, why it happened, and what happens next
Michael Roberts
2016 · Hyamarket Books
Saved 2026-07-27
9KFQCRD2
book
Capital
Karl Marx
Saved 2026-07-27
EKWDW2YM
webpage
Dejan Panovski
2018
Saved 2026-07-27
SCP copies files securely between local and remote hosts over SSH. This guide covers syntax, common options, and practical examples for everyday file transfers.
MTP7BISW
journalArticle
But Who will Monitor the Monitor?
David Rahman
Saved 2026-07-26
Consider a group of individuals in a strategic environment with moral hazard and adverse selection, and suppose that providing incentives for a given outcome requires a monitor to detect deviations. What about the monitor’s deviations? In this paper I propose a contract that makes the monitor responsible for the monitoring technology, and thereby successfully provides incentives even when the monitor’s observations are not only private, but costly, too. I also characterize exactly when such a contract can provide monitors with the right incentives to perform. In doing so, I emphasize virtual enforcement and suggest its implications for the theory of repeated games.
A8D2A4C4
journalArticle
David Rahman
2012 · American Economic Review
Saved 2026-07-26
Suppose that providing incentives for a group of individuals in a strategic context requires a monitor to detect their deviations. What about the monitor's deviations? To address this question, I propose a contract that makes the monitor responsible for monitoring, and thereby provides incentives even when the monitor's observations are not only private, but costly, too. I also characterize exactly when such a contract can provide monitors with the right incentives to perform. In doing so, I emphasize virtual enforcement and suggest its implications for the theory of repeated games. (JEL C78, D23, D82, D86)
6YYJMJVG
webpage
Saved 2026-07-26
steerable and explainable AI
VYNJUDYB
webpage
Lila Shroff, Rose Horowitch
2026
Saved 2026-07-26
AI companies are stripping universities of their best researchers.
UF2KXW3N
preprint
Jiarui Zhang, Muzi Tao, Shangshang Wang, Ollie Liu, Xuezhe Ma, Willie Neiswanger
2026
Saved 2026-07-24
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.
KW8BXXF5
forumPost
Felix Choussat
2026
Saved 2026-07-23
C8EEHYDF
journalArticle
International AI Safety Report 2026
2026
Saved 2026-07-23
ENLRWTQU
blogPost
2023
Saved 2026-07-23
Why are billions of dollars being poured into artificial intelligence R&D this year? Companies certainly expect to get a return on their investment. Arguably, the main reason AI is profitable i…
687X8XDT
webpage
Saved 2026-07-23
SECDNSGM
preprint
Joseph Carlsmith
2024
Saved 2026-07-23
This report examines what I see as the core argument for concern about existential risk from misaligned artificial intelligence. I proceed in two stages. First, I lay out a backdrop picture that informs such concern. On this picture, intelligent agency is an extremely powerful force, and creating agents much more intelligent than us is playing with fire -- especially given that if their objectives are problematic, such agents would plausibly have instrumental incentives to seek power over humans. Second, I formulate and evaluate a more specific six-premise argument that creating agents of this kind will lead to existential catastrophe by 2070. On this argument, by 2070: (1) it will become possible and financially feasible to build relevantly powerful and agentic AI systems; (2) there will be strong incentives to do so; (3) it will be much harder to build aligned (and relevantly powerful/agentic) AI systems than to build misaligned (and relevantly powerful/agentic) AI systems that are still superficially attractive to deploy; (4) some such misaligned systems will seek power over humans in high-impact ways; (5) this problem will scale to the full disempowerment of humanity; and (6) such disempowerment will constitute an existential catastrophe. I assign rough subjective credences to the premises in this argument, and I end up with an overall estimate of ~5% that an existential catastrophe of this kind will occur by 2070. (May 2022 update: since making this report public in April 2021, my estimate here has gone up, and is now at >10%.)
8V4UTSEL
webpage
Saved 2026-07-23
New research on how we've reduced agentic misalignment
DUZQSQUI
webpage
2026
Saved 2026-07-23
IUYNJRGC
journalArticle
New Generation of Counter UAS Systems to Defeat of Low Slow and Small (LSS) Air Threats
Jacco Dominicus
Saved 2026-07-22
Detecting, classifying, identifying, tracking and defeating low, slow and small air threats presents a major challenge for existing sensor and effector systems. So-called first generation Counter Unmanned Aircraft Systems (C-UAS) systems often rely on detecting the datalink from the controller to the drone which provides limited capability against current threats. However, this means of detecting drones is a challenge when operators manipulate standard datalinks and it will not work at all against current and future autonomous drones. Other current methods of detecting and neutralising drones include for example combining radar with optical sensors. These systems are not always reliable, can generate large numbers of false alerts and are often manpower intensive to operate. The NATO SCI-301 Research Task Group (RTG) has been working on specifying what second generation C-UAS systems should entail. This paper will outline the findings of this RTG over the past three years.
MPAC8FIT
forumPost
Scott Alexander
2009
Saved 2026-07-22
JF4F4CZE
forumPost
Yair Halberstadt
2026
Saved 2026-07-22
UDPJBLKL
blogPost
Aella
2025
Saved 2026-07-22
The Growing Kids God's Way protocol
K7WRRPIW
blogPost
Aella
2025
Saved 2026-07-22
and the way we treat children as property
8QQZKPHP
blogPost
Saved 2026-07-22
DUV3PX99
journalArticle
A REPORT TO THE PRESIDENT
Michael Kratsios
Saved 2026-07-22
SP2BMHRK
webpage
Saved 2026-07-22
NFY63EP3
webpage
PhD
Saved 2026-07-22
687KYGFP
journalArticle
Lukas Röseler, Leonard Kaiser, Christopher Doetsch, Noah Klett, Christian Seida, Astrid Schütz, Balazs Aczel, Nadia Adelina et al.
2024 · Journal of Open Psychology Data
Saved 2026-07-22
A9N8Z9UR
webpage
Saved 2026-07-22
Q4NETH4D
webpage
Saved 2026-07-22
52T3V9DG
preprint
Joseba Fernandez de Landa, Carla Perez-Almendros, Jose Camacho-Collados
2026
Saved 2026-07-22
LLMs have been showing limitations when it comes to cultural coverage and competence, and in some cases show regional biases such as amplifying Western and Anglocentric viewpoints. While there have been works analysing the cultural capabilities of LLMs, there has not been specific work on highlighting LLM regional preferences when it comes to cultural-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ). The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan. Moveover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs and show less inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training.
2GY4ANP8
webpage
Micah Carroll
Saved 2026-07-21
RNH8J3D3
blogPost
Alex Koren
2016
Saved 2026-07-21
I get asked a lot how to apply for the Thiel Fellowship and it usually boils down to two questions:
8X4EFBMR
webpage
Saved 2026-07-21
B46WEUWZ
preprint
Brett Reynolds
2026
Saved 2026-07-20
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. This paper introduces adversarial pragmatics as a benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. The contribution is empirical and methodological: a linguistically controlled taxonomy, an 18-item seed benchmark with validator-enforced metadata, a 54-row local seed pilot, an expert-evaluation protocol distinguishing task success, policy compliance, safety risk, refusal outcome, and evaluator confidence, and metrics for judge validity, diagnostic ambiguity, and taxonomy drift. The benchmark treats labels as inference licenses: it tests whether safety-relevant categories project across paraphrase, wrapper, model, and judge condition. In the pilot, a rubric-aided LLM judge graded its own outputs with expected-behaviour fields visible and still missed the safety-relevant minority classes.
RSHG3PFQ
preprint
Brett Reynolds
2026
Saved 2026-07-20
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. This paper introduces adversarial pragmatics as a benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. The contribution is empirical and methodological: a linguistically controlled taxonomy, an 18-item seed benchmark with validator-enforced metadata, a 54-row local seed pilot, an expert-evaluation protocol distinguishing task success, policy compliance, safety risk, refusal outcome, and evaluator confidence, and metrics for judge validity, diagnostic ambiguity, and taxonomy drift. The benchmark treats labels as inference licenses: it tests whether safety-relevant categories project across paraphrase, wrapper, model, and judge condition. In the pilot, a rubric-aided LLM judge graded its own outputs with expected-behaviour fields visible and still missed the safety-relevant minority classes.
BZ74CTB8
preprint
Otto Jespersen, Brett Reynolds, Peter Evans
2025
Saved 2026-07-20
This volume presents a new edition of Otto Jespersen's landmark 1917 study of negation in English and other languages, primarily Germanic and Romance. While best known for describing what would later be called “Jespersen's Cycle'”, this work offers far more: a comprehensive analysis of negative expressions, their forms, functions, and historical development. The book examines topics ranging from negative prefixes to the distinction between special and nexal negation, supported by Jespersen's characteristically rich collection of authentic examples. This edition features an extensive new introduction by Olli O. Silvennoinen that situates Jespersen's work in its historical and intellectual context while highlighting its continued relevance to contemporary linguistics. The main text has been entirely re-typeset to enhance readability, with examples presented in modern numbered format and Leipzig-style glosses added for non-English examples. Where possible, hyperlinks to source materials have been provided, making this classic work more accessible than ever for modern scholars and students of linguistics.
6R7X4NBG
forumPost
Zohar Atkins [@ZoharAtkins]
2026
Saved 2026-07-20
E5HTGGDS
webpage
Niall Ferguson
2026
Saved 2026-07-20
The tools that once exposed and debunked Holocaust denial are powerless against AI and the algorithm. Niall Ferguson and John-Clark Levin ask: Is there a remedy?
S7IY94V2
webpage
Saved 2026-07-20
8ZVFGMHC
webpage
Saved 2026-07-20
G7G6N7CY
webpage
2026
Saved 2026-07-20
Codex (wife) took custody of the kids (dreams and whimsy) and now i am in a social club at 1:30 am confronting my thoughts under the influence of tequila.
Z68YQZ2A
newspaperArticle
Jordi Lippe-McGraw
2026 · Wall Street Journal
Saved 2026-07-19
One nanny isn’t cutting it anymore. Some parents are spending upward of $250,000 on teams for their children. Luxury services offer potty-training, baby chefs and bike-riding lessons.
74C9HGP6
webpage
Ti Guo
Saved 2026-07-19
A collaborative AI workspace, built on your company context. Build and orchestrate agents right alongside your team's projects, meetings, and connected apps.
TPMGVC6W
journalArticle
Compulsive thalamic self-stimulation: a case with metabolic, electrophysiologic and behavioral correlates
Russell K. Portenoy, Jens O. Jarden, John J. Sidtis, Richard B. Lipton, Kathleen M. Foley, David A. Rottenberg
1986 · Pain
Saved 2026-07-19
A 48-year-old woman with a stimulating electrode implanted in the right thalamic nucleus ventralis posterolateralis developed compulsive self-stimulation associated with erotic sensations and changes in autonomic and neurologic function. Stimulation effects were evaluated by neuropsychologic testing, endocrine studies, positron emission tomographic measurements of regional cerebral metabolic rate for glucose, EEG and evoked potentials. During stimulation, vital signs and pupillary diameter increased and a left hemiparesis and left hemisensory loss developed. Verbal functions deteriorated and visuospatial processing improved. Plasma growth hormone concentrations decreased, and adrenocorticotrophic hormone and cortisol levels rose. With stimulation, glucose metabolism increased in both thalami and both hemispheres, reversing baseline right-sided hypometabolism and right-left asymmetries. EEG and both somatosensory and brain-stem auditory evoked potentials remained unchanged during stimulation, while visual evoked potentials revealed evidence of anterior visual pathway dysfunction in the left eye. This case establishes the potential for addiction to deep brain stimulation and demonstrates that widespread behavioral and physiological changes, with concomitant alteration in the regional cerebral metabolic rate for glucose, may accompany unilateral thalamic stimulation.
LBZT22M9
webpage
Saved 2026-07-19
ZIUI8XP8
journalArticle
Putting Bugs in Your Data Center Might Actually be a Good Idea
Alon Rashelbach, Mark Silberstein
Saved 2026-07-18
Data centers of cloud providers hold millions of processor cores, exabytes of storage, and petabytes of network bandwidth. Research shows that in 2019, data centers consumed more than 2% of global electricity production, where 50% of consumption targeted for cooling infrastructures. While the most effective solution for thermal distribution is liquid cooling, technical challenges and complexities make it expensive. We suggest using living spiders as cooling devices for data centers. A prior work shows that spider silk has high thermal conductivity, close to that of copper: the second-best metallic conductor. Spiders not only generate spider silk but maintain it. Recruiting spiders for the job requires no more than inserting bugs to the data center for the spiders to catch. This solution is effective, self-sustaining, and environment-friendly, but requires solving a number of non-trivial technical and zoological challenges on the way to make it practical.
MV88FNMT
preprint
Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks
2026
Saved 2026-07-17
Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). This is a limit case for behavioral oversight, where surface tokens carry no information about the underlying reasoning. But hidden from the output is not the same as hidden from us. On four task families (fact retrieval, parallel numeric composition, string manipulation, and in-context computation), two open-weights frontier models (DeepSeek V3, Kimi K2) compute over filler tokens in a structured, legible way: attention routes the question through the filler region to the answer, logit-lens readouts show retrieved facts emerging early and their composition crystallizing in late layers, and KV-cache transplants at filler positions causally swap outputs between examples. We introduce an unsupervised decoding pipeline that takes only hidden states as input and recovers intermediate values with 80-95% accuracy (best LLM judge) across both models and all four tasks, without ground-truth labels or training. Hidden computation that defeats behavioral CoT monitoring is, on these tasks, directly readable from the residual stream, suggesting monitorability is a property of the model's full computational trace, not just its surface tokens.
RRKVN3S5
journalArticle
Taifeng Liu, Xin Zhou, Michel Dupuis, Can Li
2015 · Physical Chemistry Chemical Physics
Saved 2026-07-17
T6W8ZKV6
webpage
Rajan Agarwal
Saved 2026-07-17
Lessons learned from building multiplayer world models. Built a video tokenizer with spatial attention and a dynamics model with action spaces.
J6A5SY55
blogPost
Matteo Cococcioni
2023
Saved 2026-07-16
28 August – 1 September 2023 University of Pavia (Italy) From August 28th to September 1st, 2023, Pavia (Italy) hosted the first in-person edition of the “Advanced Quantum ESPRESSO Scho…
8UL6I8JJ
forumPost
Zvi
2025
Saved 2026-07-16
DDKAL79Z
forumPost
Johann Kurtz [@JohannKurtz]
2026
Saved 2026-07-16
7NGQ6VIU
journalArticle
Saiying Steenbergen-Hu, Matthew C. Makel, Paula Olszewski-Kubilius
2016 · Review of Educational Research
Saved 2026-07-16
QRYVLDA8
journalArticle
A. Floris, I. Timrov, B. Himmetoglu, N. Marzari, S. De Gironcoli, M. Cococcioni
2020 · Physical Review B
Saved 2026-07-16
VUEQEXJY
forumPost
Joseph Miller
2026
Saved 2026-07-15
GBG6IPCP
journalArticle
Taifeng Liu, Qianyu Zhao, Can Li, Yang Lyu, Al Et.
2019 · The Journal of Physical Chemistry C
Saved 2026-07-14
Q86MIVXM
journalArticle
Taifeng Liu, Mengsi Cui, Michel Dupuis
2020 · The Journal of Physical Chemistry C
Saved 2026-07-14
3NE7BD8B
journalArticle
Julia Wiktor, Francesco Ambrosio, Alfredo Pasquarello
2018 · ACS Energy Letters
Saved 2026-07-14
QHNZDWB8
journalArticle
Francesco Ambrosio, Julia Wiktor, Alfredo Pasquarello
2018 · ACS Applied Materials & Interfaces
Saved 2026-07-14
W2T9CEYY
journalArticle
Julia Wiktor, Alfredo Pasquarello
2019 · ACS Applied Materials & Interfaces
Saved 2026-07-14
QEALA5LS
journalArticle
John P. Perdew, Kieron Burke, Matthias Ernzerhof
1996 · Physical Review Letters · American Physical Society
Saved 2026-07-14
Generalized gradient approximations (GGA's) for the exchange-correlation energy improve upon the local spin density (LSD) description of atoms, molecules, and solids. We present a simple derivation of a simple GGA, in which all parameters (other than those in LSD) are fundamental constants. Only general features of the detailed construction underlying the Perdew-Wang 1991 (PW91) GGA are invoked. Improvements over PW91 include an accurate description of the linear response of the uniform electron gas, correct behavior under uniform scaling, and a smoother potential.
G36NYYKD
journalArticle
Graeme Henkelman, Andri Arnaldsson, Hannes Jónsson
2006 · Computational Materials Science
Saved 2026-07-13
An algorithm is presented for carrying out decomposition of electronic charge density into atomic contributions. As suggested by Bader [R. Bader, Atoms in Molecules: A Quantum Theory, Oxford University Press, New York, 1990], space is divided up into atomic regions where the dividing surfaces are at a minimum in the charge density, i.e. the gradient of the charge density is zero along the surface normal. Instead of explicitly finding and representing the dividing surfaces, which is a challenging task, our algorithm assigns each point on a regular (x,y,z) grid to one of the regions by following a steepest ascent path on the grid. The computational work required to analyze a given charge density grid is approximately 50 arithmetic operations per grid point. The work scales linearly with the number of grid points and is essentially independent of the number of atoms in the system. The algorithm is robust and insensitive to the topology of molecular bonding. In addition to two test problems involving a water molecule and NaCl crystal, the algorithm has been used to estimate the electrical activity of a cluster of boron atoms in a silicon crystal. The highly stable three-atom boron cluster, B3I is found to have a charge of −1.5e, which suggests approximately 50% reduction in electrical activity as compared with three substitutional boron atoms.
BB7TBSSP
book
Atoms in Molecules: A Quantum Theory
Richard F. W. Bader
1994 · Oxford University Press
Saved 2026-07-13
The molecular structure hypothesis--that a molecule is a collection of atoms linked by a network of bonds-- provides the principal means of ordering and classifying observations in chemistry. However this hypothesis is not related directly to the physics which governs the motions of atomic nuclei and electrons. It is the purpose of this important new book to show that a theory can be developed to establish the molecular structure hypothesis, demonstrating that the atoms in a molecule are real, with properties predicted and defined by the laws of quantum mechanics, and that the structure their presence imparts to a molecule is indeed a consequence of the underlying physics. As a result, the classification based upon the concept of atoms in molecules is freed from its empirical constraints and the full predictive power of quantum mechanics can be incorporated into the resulting theory--a theory of atoms in molecules. Eminently accessible and readable, the book will interest all scientists involved with experiment and observation at the atomic level, in addition to theoreticians. , The molecular structure hypothesis--that a molecule is a collection of atoms linked by a network of bonds-- provides the principal means of ordering and classifying observations in chemistry. However this hypothesis is not related directly to the physics which governs the motions of atomic nuclei and electrons. It is the purpose of this important new book to show that a theory can be developed to establish the molecular structure hypothesis, demonstrating that the atoms in a molecule are real, with properties predicted and defined by the laws of quantum mechanics, and that the structure their presence imparts to a molecule is indeed a consequence of the underlying physics. As a result, the classification based upon the concept of atoms in molecules is freed from its empirical constraints and the full predictive power of quantum mechanics can be incorporated into the resulting theory--a theory of atoms in molecules. Eminently accessible and readable, the book will interest all scientists involved with experiment and observation at the atomic level, in addition to theoreticians.
DRKP2N4F
preprint
Ilyes Batatia, Dávid Péter Kovács, Gregor N. C. Simm, Christoph Ortner, Gábor Csányi
2023
Saved 2026-07-13
Creating fast and accurate force fields is a long-standing challenge in computational chemistry and materials science. Recently, several equivariant message passing neural networks (MPNNs) have been shown to outperform models built using other approaches in terms of accuracy. However, most MPNNs suffer from high computational cost and poor scalability. We propose that these limitations arise because MPNNs only pass two-body messages leading to a direct relationship between the number of layers and the expressivity of the network. In this work, we introduce MACE, a new equivariant MPNN model that uses higher body order messages. In particular, we show that using four-body messages reduces the required number of message passing iterations to just two, resulting in a fast and highly parallelizable model, reaching or exceeding state-of-the-art accuracy on the rMD17, 3BPA, and AcAc benchmark tasks. We also demonstrate that using higher order messages leads to an improved steepness of the learning curves.
VWADBXMR
journalArticle
Ravishankar Sundararaman, William A., III Goddard
2015 · The Journal of Chemical Physics
Saved 2026-07-13
Many important applications of electronic structure methods involve molecules or solid surfaces in a solvent medium. Since explicit treatment of the solvent in such methods is usually not practical, calculations often employ continuum solvation models to approximate the effect of the solvent. Previous solvation models either involve a parametrization based on atomic radii, which limits the class of applicable solutes, or based on solute electron density, which is more general but less accurate, especially for charged systems. We develop an accurate and general solvation model that includes a cavity that is a nonlocal functional of both solute electron density and potential, local dielectric response on this nonlocally determined cavity, and nonlocal approximations to the cavity-formation and dispersion energies. The dependence of the cavity on the solute potential enables an explicit treatment of the solvent charge asymmetry. With four parameters per solvent, this “CANDLE” model simultaneously reproduces solvation energies of large datasets of neutral molecules, cations, and anions with a mean absolute error of 1.8 kcal/mol in water and 3.0 kcal/mol in acetonitrile.
XASNDNSB
journalArticle
Oliviero Andreussi, Ismaila Dabo, Nicola Marzari
2012 · The Journal of Chemical Physics
Saved 2026-07-13
The solvation model proposed by Fattebert and Gygi [Journal of Computational Chemistry 23, 662 (2002)] and Scherlis et al. [Journal of Chemical Physics 124, 074103 (2006)] is reformulated, overcoming some of the numerical limitations encountered and extending its range of applicability. We first recast the problem in terms of induced polarization charges that act as a direct mapping of the self-consistent continuum dielectric; this allows to define a functional form for the dielectric that is well behaved both in the high-density region of the nuclear charges and in the low-density region where the electronic wavefunctions decay into the solvent. Second, we outline an iterative procedure to solve the Poisson equation for the quantum fragment embedded in the solvent that does not require multi-grid algorithms, is trivially parallel, and can be applied to any Bravais crystallographic system. Last, we capture some of the non-electrostatic or cavitation terms via a combined use of the quantum volume and quantum surface [Physical Review Letters 94, 145501 (2005)] of the solute. The resulting self-consistent continuum solvation (SCCS) model provides a very effective and compact fit of computational and experimental data, whereby the static dielectric constant of the solvent and one parameter allow to fit the electrostatic energy provided by the PCM model with a mean absolute error of 0.3 kcal/mol on a set of 240 neutral solutes. Two parameters allow to fit experimental solvation energies on the same set with a mean absolute error of 1.3 kcal/mol. A detailed analysis of these results, broken down along different classes of chemical compounds, shows that several classes of organic compounds display very high accuracy, with solvation energies in error of 0.3-0.4 kcal/mol, whereby larger discrepancies are mostly limited to self-dissociating species and strong hydrogen-bond forming compounds.
6AC7HIDE
journalArticle
P Giannozzi, O Andreussi, T Brumme, O Bunau, M Buongiorno Nardelli, M Calandra, R Car, C Cavazzoni et al.
2017 · Journal of Physics: Condensed Matter · IOP Publishing
Saved 2026-07-13
Quantum ESPRESSO is an integrated suite of open-source computer codes for quantum simulations of materials using state-of-the-art electronic-structure techniques, based on density-functional theory, density-functional perturbation theory, and many-body perturbation theory, within the plane-wave pseudopotential and projector-augmented-wave approaches. Quantum ESPRESSO owes its popularity to the wide variety of properties and processes it allows to simulate, to its performance on an increasingly broad array of hardware architectures, and to a community of researchers that rely on its capabilities as a core open-source development platform to implement their ideas. In this paper we describe recent extensions and improvements, covering new methodologies and property calculators, improved parallelization, code modularization, and extended interoperability both within the distribution and with external software.
BKUW28EG
journalArticle
Á. Valdés, Z.-W. Qu, G.-J. Kroes, J. Rossmeisl, J. K. Nørskov
2008 · The Journal of Physical Chemistry C
Saved 2026-07-13
3TZ4KEIB
conferencePaper
Mingyang Xu, Yanheng Li, Burcu Nimet Dumlu, RAY LC, Giulia Barbareschi, Matthias Hoppe, Jie Li, Kouta Minamizawa et al.
2026 · Proceedings of the 2026 Designing Interactive Systems Conference · Association for Computing Machinery
Saved 2026-07-13
Soft floating robots (SFRs) represent a shift from rigid machines, offering gravity-defying, compliant, and tactile embodiments for indoor cohabitation. However, their development remains fragmented across isolated prototypes, lacking a coherent design vocabulary. Without a systematic understanding of their interactional capabilities, designers struggle to leverage SFRs’ unique affordances, and these systems often remain limited to novelty applications that are difficult to integrate into everyday life. To address this, we propose a design space for interaction with SFRs. Informed by an exploratory study with 12 experts from HCI, Design, and Robotics, we identify ten design dimensions spanning physical, interactive, and behavioral properties, along with a range of application scenarios. We further present proof-of-concept design examples to demonstrate how this design space can support diverse interaction possibilities. This work contributes a structured framework for understanding and designing interactions with SFRs, supporting their integration into everyday indoor environments.
SPTVU4NG
preprint
Yujing Wei, John L. Weber, James M. Stevenson, Zachary K. Goldsmith, Xiaowei Xie, Leif D. Jacobson, Richard A. Friesner
2026
Saved 2026-07-13
Machine learning interatomic potentials (MLIPs), also known as machine learning force fields (MLFFs), offer scalable means of simulating complex systems and processes at \textit{ab initio} level accuracy. One such process is the critical yet still poorly understood formation of the solid electrolyte interphase (SEI) at the anode of a Li-ion battery (LIB) during the first charge cycle, where electrochemical reduction of the electrolyte leads to the generation of decomposition products. MLIPs are uniquely poised to atomistically describe these electrochemical processes, as they are not as affected by the same limitations in bonding and electron transfer as classical force fields. Nonetheless, training MLIPs to run accurate dynamics of a condensed phase with two different oxidation states, such as in electrochemistry, is challenging for many architectures. In this work, we show that by using MPNICE, a message passing MLIP architecture with iterative charge equilibration, we are able to accurately (within 1 kcal/mol) train models along two potential energy surfaces (reduced and unreduced) for LIB-relevant electrolyte systems. Importantly, we demonstrate strategies for sampling and training to examples of anion radicals of these species, which often are not centered on any atom (off-center radicals, or OCRs). We additionally discuss well known limitations of global charge equilibration (Qeq) algorithms in erroneously de-localizing charge, and test methods to alleviate the impact on resulting dynamics. Simulations using these models reveal new insights into electrolyte reduction and considerations for the realistic simulation of electron transfer processes in the condensed phase.
FQBRRQQR
journalArticle
Zeyu Wang, William A. Goddard, Hai Xiao
2023 · Nature Communications · Nature Publishing Group
Saved 2026-07-13
Oxygen evolution reaction (OER) is of crucial importance to sustainable energy and environmental engineering, and layered double hydroxides (LDHs) are among the most active catalysts for OER in alkaline conditions, but the reaction mechanism for OER on LDHs remains controversial. Distinctive types of reaction mechanisms have been proposed for the O-O coupling in OER, yet they compose a coupled reaction network with competing kinetics dependent on applied potentials. Herein, we combine grand-canonical methods and micro-kinetic modeling to unravel that the nature of dominant mechanism for OER on LDHs transitions among distinctive types as a function of applied potential, and this arises from the interplay among applied potential and competing kinetics in the coupled reaction network. The theory-predicted overpotentials, Tafel slopes, and findings are in agreement with the observations of experiments including isotope labelling. Thus, we establish a computational methodology to identify and elucidate the potential-dependent mechanisms for electrochemical reactions.
8IYE7RY5
journalArticle
Zhi Wei Seh, Jakob Kibsgaard, Colin F. Dickens, Ib Chorkendorff, Jens K. Nørskov, Thomas F. Jaramillo
2017 · Science
Saved 2026-07-13
Electrocatalysis plays a central role in clean energy conversion, enabling a number of sustainable processes for future technologies. This review discusses design strategies for state-of-the-art heterogeneous electrocatalysts and associated materials for several different electrochemical transformations involving water, hydrogen, and oxygen, using theory as a means to rationalize catalyst performance. By examining the common principles that govern catalysis for different electrochemical reactions, we describe a systematic framework that helps to understand trends in catalyzing these reactions, serving as a guide to new catalyst development, while highlighting key gaps that need to be addressed. We conclude by extending this framework to emerging clean energy reactions including hydrogen peroxide production, carbon dioxide reduction and nitrogen reduction, where the development of improved catalysts could allow for the sustainable production of a broad range of fuels and chemicals.
QXI35JPB
journalArticle
Chang Liu, Jin Qian, Yifan Ye, Hua Zhou, Cheng-Jun Sun, Colton Sheehan, Zhiyong Zhang, Gang Wan et al.
2021 · Nature Catalysis · Nature Publishing Group
Saved 2026-07-13
Efficient electrocatalysts for the oxygen evolution reaction (OER) are paramount to the development of electrochemical devices for clean energy and fuel conversion. However, the structural complexity of heterogeneous electrocatalysts makes it a great challenge to elucidate the surface catalytic sites and OER mechanisms. Here, we report that catalytic single-site Co in a well-defined brookite TiO2 nanorod (210) surface (Co-TiO2) presents turnover frequencies that are among the highest for Co-based heterogeneous catalysts reported to date, reaching 6.6 ± 1.2 and 181.4 ± 28 s−1 at 300 and 400 mV overpotentials, respectively. Based on grand canonical quantum mechanics calculations and the single-site Co atomic structure validated by in situ and ex situ spectroscopic probes, we have established a full description of the catalytic reaction kinetics for Co-TiO2 as a function of applied potential, revealing an adsorbate evolution mechanism for the OER. The computationally predicted Tafel slope and turnover frequencies exhibit exceedingly good agreement with experiment.
YSB43FIB
journalArticle
Chang Liu, Jin Qian, Yifan Ye, Hua Zhou, Cheng-Jun Sun, Colton Sheehan, Zhiyong Zhang, Gang Wan et al.
2022 · Nature Catalysis · Nature Publishing Group
Saved 2026-07-13
JUKGU79E
journalArticle
Chang Liu, Jin Qian, Yifan Ye, Hua Zhou, Cheng-Jun Sun, Colton Sheehan, Zhiyong Zhang, Gang Wan et al.
2021 · Nature Catalysis · Nature Publishing Group
Saved 2026-07-13
Efficient electrocatalysts for the oxygen evolution reaction (OER) are paramount to the development of electrochemical devices for clean energy and fuel conversion. However, the structural complexity of heterogeneous electrocatalysts makes it a great challenge to elucidate the surface catalytic sites and OER mechanisms. Here, we report that catalytic single-site Co in a well-defined brookite TiO2 nanorod (210) surface (Co-TiO2) presents turnover frequencies that are among the highest for Co-based heterogeneous catalysts reported to date, reaching 6.6 ± 1.2 and 181.4 ± 28 s−1 at 300 and 400 mV overpotentials, respectively. Based on grand canonical quantum mechanics calculations and the single-site Co atomic structure validated by in situ and ex situ spectroscopic probes, we have established a full description of the catalytic reaction kinetics for Co-TiO2 as a function of applied potential, revealing an adsorbate evolution mechanism for the OER. The computationally predicted Tafel slope and turnover frequencies exhibit exceedingly good agreement with experiment.
EWIE24UH
journalArticle
Chang Liu, Soonho Kwon, Perrin Godbold, Grayson Johnson, Sooyeon Hwang, Chengjun Sun, Hua Zhou, William A. Goddard et al.
2025 · Journal of the American Chemical Society
Saved 2026-07-13
The design of advanced electrocatalysts is often hindered by uncertainties in identifying and controlling the active surfaces and catalytic centers within heterogeneous materials. Here we present the synthesis of single-site Co catalysts, substitutionally doped into surface-controlled TiO2 anatase nanocrystals, aimed at enhancing the oxygen evolution reaction (OER). Grand canonical quantum mechanics calculations reveal that the kinetics of the OER, following an adsorbate evolution mechanism, is markedly influenced by the coordination environment of Co. The simulations suggest significantly higher turnover frequencies when Co is doped into the (001) surface of TiO2 compared to the (101) surface. Consistent with the computational findings, experimental results show that Co-doped TiO2 (Co-TiO2) nanoplates with selectively exposed {001} surfaces exhibit enhanced current densities and turnover frequencies compared to Co-TiO2 nanobipyramids with {101} surfaces. This study highlights the synergy between theoretical calculations and precision synthesis in the development of more effective catalysts.
TUB4ACFF
journalArticle
Okan K. Orhan, David D. O'Regan
2020 · Physical Review B · American Physical Society
Saved 2026-07-13
Titanium dioxide (TiO2) presents a long-standing challenge for approximate Kohn-Sham density functional theory (KS-DFT), as well as to its Hubbard-corrected extension, DFT+U. We find that a previously proposed extension of first-principles DFT+U to incorporate a Hund's 𝐽 correction, termed DFT+U+J, in combination with parameters calculated using a recently proposed linear-response theory, predicts fundamental band gaps that are accurate to well within the experimental uncertainty in rutile and anatase TiO2. Our approach builds upon established findings that Hubbard correction of both the titanium 3⁢𝑑 and oxygen 2⁢𝑝 subspaces in TiO2, symbolically giving DFT+U𝑑,𝑝, is necessary to achieve acceptable band gaps using DFT+U. This requirement remains when the first-principles Hund's 𝐽 is included. We also find that the calculated gap depends on the correlated subspace definition even when using subspace-specific first-principles 𝑈 and 𝐽 parameters. Using the simplest reasonable correlated subspace definition and underlying functional, the local density approximation, we show that high accuracy results from using a relatively uncomplicated form of the DFT+U+J functional. For closed-shell systems such as TiO2, we describe how various DFT+U+J functionals reduce to DFT+U with suitably modified parameters, so that reliable band gaps can be calculated for rutile and anatase with no modifications to a conventional DFT+U code.
I5WK9K9M
preprint
Christian S. Ahart, Denan Li, Jochen Blumberger, Shi Liu
2026
Saved 2026-07-13
Transition metal oxides have attracted much attention as photo(electrochemical)-catalysts but practical applications are typically hampered by their low and anisotropic charge mobility. A deep understanding of excess charge carrier transport in these materials requires a dynamical treatment of nuclear motion that goes well beyond standard approaches. Here we introduce DeepPolaron, a machine learning framework boosting the accessible time scale of first principles molecular dynamics of adiabatic polaron transport by three orders of magnitude at a virtually negligible loss in accuracy. We apply our method to excess electron and hole transport in titanium dioxide rutile and anatase. We find that the excess electron in rutile relaxes to a polaron predominantly localized on a single Ti atom with hopping occurring only along the [001] direction, associated with an activation energy of 39 meV and a room temperature mobility of $4.4 \times 10^{-2}$ cm$^2$/Vs in good agreement with experiment. In contrast the hole polaron in anatase is localized on a single O atom, and due to poor O 2p orbital overlap with first nearest neighbors charge transport occurs primarily to second nearest neighbors, with a large activation energy of 139 meV resulting in a small room temperature mobility of $1.4 \times 10^{-3}$ cm$^2$/Vs. This work provides a finite temperature first-principles characterization of small polaron transport in rutile and anatase, with a methodology that is directly transferable to other small polaron forming materials and interfacial charge-transfer processes.
RAZ7C5VB
journalArticle
Xiaodan Yan, Xiao Han, Jinlu He
2025 · The Journal of Physical Chemistry Letters · American Chemical Society
Saved 2026-07-13
Using time-dependent density functional theory (TD-DFT) and nonadiabatic molecular dynamics (NAMD) simulations, we elucidate how hydrogen bonding and water dynamics regulate hole transfer at the anatase TiO2(101)/water interface. Compared to low-density water (LW), moderate-density water (MW) enhances hydrogen bonding between surface- and nonsurface-adsorbed water, restricting interfacial water mobility. This suppresses nonadiabatic coupling and slows the hole transfer. Conversely, higher-density water (HW) near the vacuum destabilizes hydrogen bonding networks, freeing interfacial water and amplifying thermal motion. Enhanced disorder strengthens nonadiabatic coupling, accelerating the hole transfer. Temperature-dependent simulations show that thermal energy overcomes hydrogen bonding constraints: elevated temperatures intensify water dynamics and nonadiabatic coupling, accelerating hole transfer. These results establish hydrogen bonding and thermal fluctuations as key regulators of charge dynamics at the semiconductor/liquid interfaces.
2QJRHC5L
journalArticle
Chaitanya B. Hiragond, Prabhat Prakash, Tridip Das, Niket S. Powar, Eunhee Gong, Jeonghyeon Lee, Jin-Woo Jung, Chang-Hee Cho et al.
2026 · Applied Catalysis B: Environment and Energy
Saved 2026-07-13
Photocatalytic CO2 reduction to value-added chemicals is limited by inefficient charge transfer and sluggish multielectron kinetics. Here, we develop a TiO2 (P25) based ternary system incorporating Pt NPs and 1T-dominant MoSe2 NSs as dual cocatalysts (i.e., Pt/TiO2-MoSe2) to direct charge flow and reaction pathways. This Pt/TM architecture accelerates charge separation and provides highly active sites for CO2 activation. The optimized Pt1.5%-TiO2-MoSe2 exhibits a CH4 evolution rate of 17.81 μmol g−1 under 5-sun illumination with ≈ 98% selectivity in the gas phase, achieving a 65-fold enhancement over TiO2 (P25). Under multi-sun irradiation, increased photon flux boosts activity, indicating the critical role of charge carrier density, and facilitates rapid CH4 desorption, thereby overcoming the activity-selectivity trade-off. The coexistence of static interfacial charge redistribution and dynamic photoinduced electron transfer, validated by experiment (XPS, XAS) and Density Functional Theory (DFT) quantum mechanics (QM) calculations, ensures the retention of the active 1T-MoSe2 phase and establishes an efficient TiO2 → MoSe2 → Pt charge-funneling highway. In situ DRIFTS, combined with simulated infrared spectra (DFT), identifies a *CHO-dominated pathway for the conversion (*CO→ *CHO → *CHOH → *CH2OH → *CH2 → *CH3 → CH4), and isotopic labeling confirms the CO2-to-CH4 formation. DFT results further support the cascade charge transfer and reduced energy barriers enabled by the dual cocatalyst system.
BJUGD5LT
journalArticle
Marcos F. Calegari Andrade, Hsin-Yu Ko, Linfeng Zhang, Roberto Car, Annabella Selloni
2020 · Chemical Science
Saved 2026-07-13
TiO 2 is a widely used photocatalyst in science and technology and its interface with water is important in fields ranging from geochemistry to biomedicine. , TiO 2 is a widely used photocatalyst in science and technology and its interface with water is important in fields ranging from geochemistry to biomedicine. Yet, it is still unclear whether water adsorbs in molecular or dissociated form on TiO 2 even for the case of well-defined crystalline surfaces. To address this issue, we simulated the TiO 2 –water interface using molecular dynamics with an ab initio -based deep neural network potential. Our simulations show a dynamical equilibrium of molecular and dissociative adsorption of water on TiO 2 . Water dissociates through a solvent-assisted concerted proton transfer to form a pair of short-lived hydroxyl groups on the TiO 2 surface. Molecular adsorption of water is Δ F = 8.0 ± 0.9 kJ mol −1 lower in free energy than the dissociative adsorption, giving rise to a 5.6 ± 0.5% equilibrium water dissociation fraction at room temperature. Due to the relevance of surface hydroxyl groups to the surface chemistry of TiO 2 , our model might be key to understanding phenomena ranging from surface functionalization to photocatalytic mechanisms.
QAP33YME
journalArticle
Zezhu Zeng, Felix Wodaczek, Keyang Liu, Frederick Stein, Jürg Hutter, Ji Chen, Bingqing Cheng
2023 · Nature Communications · Nature Publishing Group
Saved 2026-07-13
Water adsorption and dissociation processes on pristine low-index TiO2 interfaces are important but poorly understood outside the well-studied anatase (101) and rutile (110). To understand these, we construct three sets of machine learning potentials that are simultaneously applicable to various TiO2 surfaces, based on three density-functional-theory approximations. Here we show the water dissociation free energies on seven pristine TiO2 surfaces, and predict that anatase (100), anatase (110), rutile (001), and rutile (011) favor water dissociation, anatase (101) and rutile (100) have mostly molecular adsorption, while the simulations of rutile (110) sensitively depend on the slab thickness and molecular adsorption is preferred with thick slabs. Moreover, using an automated algorithm, we reveal that these surfaces follow different types of atomistic mechanisms for proton transfer and water dissociation: one-step, two-step, or both. These mechanisms can be rationalized based on the arrangements of water molecules on the different surfaces. Our finding thus demonstrates that the different pristine TiO2 surfaces react with water in distinct ways, and cannot be represented using just the low-energy anatase (101) and rutile (110) surfaces.
8RAYIWE4
journalArticle
Chuin Wei Tan, Marc L. Descoteaux, Mit Kotak, Gabriel De Miranda Nascimento, Seán R. Kavanagh, Laura Zichi, Menghang Wang, Aadit Saluja et al.
2026 · Digital Discovery
Saved 2026-07-13
The NequIP framework is redesigned for scalable distributed training and PyTorch 2.0 compilation. AOT Inductor inference and optimized Allegro kernels accelerate molecular dynamics by factors of 5–18 on practical system sizes. , Machine learning interatomic potentials, particularly those based on deep equivariant neural networks, have demonstrated state-of-the-art accuracy and computational efficiency in atomistic modeling tasks like molecular dynamics and high-throughput screening. The size of datasets and demands of downstream workflows are growing rapidly, making robust and scalable software essential. This work presents a major overhaul of the NequIP framework focusing on multi-node parallelism, computational performance, and extensibility. The redesigned framework supports distributed training on large datasets and removes barriers preventing full utilization of the PyTorch 2.0 compiler at train time. We demonstrate this acceleration in a case study by training Allegro models on the SPICE 2 dataset of organic molecular systems. For inference, we introduce the first end-to-end infrastructure that uses the PyTorch Ahead-of-Time Inductor compiler for machine learning interatomic potentials. Additionally, we implement a custom kernel for the Allegro model's most expensive operation, the tensor product. Together, these advancements speed up molecular dynamics calculations on system sizes of practical relevance by up to factors of 5 to 18.
UQHTS95T
journalArticle
Jörg Behler, Michele Parrinello
2007 · Physical Review Letters · American Physical Society
Saved 2026-07-13
The accurate description of chemical processes often requires the use of computationally demanding methods like density-functional theory (DFT), making long simulations of large systems unfeasible. In this Letter we introduce a new kind of neural-network representation of DFT potential-energy surfaces, which provides the energy and forces as a function of all atomic positions in systems of arbitrary size and is several orders of magnitude faster than DFT. The high accuracy of the method is demonstrated for bulk silicon and compared with empirical potentials and DFT. The method is general and can be applied to all types of periodic and nonperiodic systems.
ITHDGG3K
journalArticle
Akira Fujishima, Kenichi Honda
1972 · Nature · Nature Publishing Group
Saved 2026-07-13
ALTHOUGH the possibility of water photolysis has been investigated by many workers, a useful method has only now been developed. Because water is transparent to visible light it cannot be decomposed directly, but only by radiation with wavelengths shorter than 190 nm (ref. 1).
L5EARR88
webpage
GradHacker
Saved 2026-07-13
How to use a research internship to prepare for grad school.
Z82Q4DSC
blogPost
Evan Peck
2018
Saved 2026-07-13
Inclusive, student-centered research environments that prioritize mentorship and development need new models and new guidelines.
UG5QPV3V
blogPost
Emma Lurie
2021
Saved 2026-07-13
This post demystifies some of the expectations and interactions important to success in undergraduate research.
MI7YJ4LR
webpage
Saved 2026-07-13
TRDM3XMB
webpage
Saved 2026-07-13
MUSZDCUL
blogPost
Danielle Strachman, areoform
2024
Saved 2026-07-12
1517's Teen Fall Camp is back! So back!
536ERY36
webpage
Kaitlyn Tiffany
2026
Saved 2026-07-11
An evolutionary psychologist is challenging the popular understanding of kids and technology.
X9ZN5SRB
webpage
Jean M. Twenge
2023
Saved 2026-07-11
New analyses from Europe, Japan, and South Korea show increases that in many cases began before COVID.
3QX4ZXJE
webpage
Anna Gát
2021
Saved 2026-07-11
What is in your head.
7HKSREFF
blogPost
Michael Roberts
2026
Saved 2026-07-10
Part One: the revolution On the 4th July 2026 it will be 250 years since the 13 British colonies in North America declared independence from the British colonial power at a congress in Philadelphia…
HDL8X76X
webpage
Saved 2026-07-09
A detailed forecast and recommendation for how the US, China and the rest of the world should navigate superintelligence.
FB6UIW5W
blogPost
Alex Tabarrok
2026
Saved 2026-07-08
Spider Noir (Prime): I’ve had enough of the Marvel multiverse so I was worried about Spider-Noir. The writers, however, have written an excellent noir in the style of Raymond Chandler with Nicholas Cage channeling Humphrey Bogart. The Spiderman stuff is all there but it is appropriately embedded. There are some excellent lines. Most notably an […]
PX5GNXBR
forumPost
Eliezer Yudkowsky
2008
Saved 2026-07-07
BM5PRQC7
webpage
2022
Saved 2026-07-07
thinking about scary things • examples from Wave • examples from elsewhere • finding a buddy • getting the timing right • a list of abyss questions
CUYIYHWX
webpage
Saved 2026-07-07
9FEZR4TE
journalArticle
Notes on Diffy Qs: Differential Equations for Engineers
Jiri Lebl
Saved 2026-06-28
M57VR6M2
webpage
Need a VectorDB for Your GenAI Apps?Zilliz Cloud is a managed vector database built on Milvus perfect for building GenAI applications Try Free
Saved 2026-06-28
Mean pooling is commonly used to create sentence embeddings from transformer token outputs like BERT because it provides
3ZC4AYWE
webpage
Brian Potter
2024
Saved 2026-06-27
Electric power in the US is provided by the electrical grid, a huge network of power plants, transmission lines, and transformers that moves electric power from where it's generated to where it's consumed.
225YTHCC
preprint
Zhe Yin, Xiaodong Gu, Beijun Shen
2026
Saved 2026-06-27
Code language models excel on code intelligence tasks, yet their internal interpretability is underexplored. Existing neuron interpretability techniques from NLP are suboptimal for source code due to programming languages formal, hierarchical, and executable nature. We empirically investigate code LLMs at the neuron level, localizing language-specific neurons (selectively responsive to one language) and concept layers (feed-forward layers encoding language-agnostic code representations). We analyze Llama-3.1-8B and Qwen2.5-Coder-32B on multilingual inputs in C++, Java, Python, Go, and JavaScript, measuring neuron selectivity and layerwise contributions during generation. We find (1) neurons specialized for individual languages alongside a universal subset supporting general-purpose generation; and (2) lower layers mainly encode language-specific syntax, while middle layers capture semantic abstractions shared across languages, emerging as concept layers. We demonstrate utility on three tasks: neuron-guided fine-tuning for code generation, clone detection via concept-layer embeddings, and concept-layer-guided transfer for code summarization, each yielding consistent gains in multilingual settings.
DD4WMCVR
journalArticle
Guruprasad Raghavan, Bahey Tharwat, Surya Narayanan Hari, Dhruvil Satani, Rex Liu, Matt Thomson
2024 · Nature Machine Intelligence · Nature Publishing Group
Saved 2026-06-27
Contemporary machine learning algorithms train artificial neural networks by setting network weights to a single optimized configuration through gradient descent on task-specific training data. The resulting networks can achieve human-level performance on natural language processing, image analysis and agent-based tasks, but lack the flexibility and robustness characteristic of human intelligence. Here we introduce a differential geometry framework—functionally invariant paths—that provides flexible and continuous adaptation of trained neural networks so that secondary tasks can be achieved beyond the main machine learning goal, including increased network sparsification and adversarial robustness. We formulate the weight space of a neural network as a curved Riemannian manifold equipped with a metric tensor whose spectrum defines low-rank subspaces in weight space that accommodate network adaptation without loss of prior knowledge. We formalize adaptation as movement along a geodesic path in weight space while searching for networks that accommodate secondary objectives. With modest computational resources, the functionally invariant path algorithm achieves performance comparable with or exceeding state-of-the-art methods including low-rank adaptation on continual learning, sparsification and adversarial robustness tasks for large language models (bidirectional encoder representations from transformers), vision transformers (ViT and DeIT) and convolutional neural networks.
RLS6SVT4
journalArticle
Exploring non-invasive sexing of early chick embryos in intact eggs using Laser Speckle Contrast Imaging (LSCI) and Deep Neural Network (DNN)
Simon Mahler, Anika Arora, Carol Readhead, Siyuan Yin, Surya Narayanan Hari, Ellie Wang, Cecilia I Moxley, Abdullahi A Adeboye et al.
Saved 2026-06-27
JY4A4UX3
journalArticle
Jackson Nyman, Thomas Denize, Ziad Bakouny, Chris Labaki, Breanna M. Titchen, Kevin Bi, Surya Narayanan Hari, Jacob Rosenthal et al.
2023 · Cell Reports Medicine
Saved 2026-06-27
ANIMUGTR
journalArticle
Jacob Rosenthal, Ryan Carelli, Mohamed Omar, David Brundage, Ella Halbert, Jackson Nyman, Surya N. Hari, Eliezer M. Van Allen et al.
2022 · Molecular Cancer Research
Saved 2026-06-27
Abstract Imaging datasets in cancer research are growing exponentially in both quantity and information density. These massive datasets may enable derivation of insights for cancer research and clinical care, but only if researchers are equipped with the tools to leverage advanced computational analysis approaches such as machine learning and artificial intelligence. In this work, we highlight three themes to guide development of such computational tools: scalability, standardization, and ease of use. We then apply these principles to develop PathML, a general-purpose research toolkit for computational pathology. We describe the design of the PathML framework and demonstrate applications in diverse use cases. PathML is publicly available at www.pathml.com.
5F8FLJ28
preprint
Jialong Jiang, David A. Sivak, Matt Thomson
2019
Saved 2026-06-27
The inverse statistical problem of finding direct interactions in complex networks is difficult. In the natural sciences, well-controlled perturbation experiments are widely used to probe the structure of complex networks. However, our understanding of how and why perturbations aid inference remains heuristic, and we lack automated procedures that determine network structure by combining inference and perturbation. Therefore, we propose a general mathematical framework to study inference with iteratively applied perturbations. Using the formulation of information geometry, our framework quantifies the difficulty of inference and the information gain from perturbations through the curvature of the underlying parameter manifold, measured by Fisher information. We apply the framework to the inference of spin network models and find that designed perturbations can reduce the sampling complexity by $10^6$-fold across a variety of network architectures. Physically, our framework reveals that perturbations boost inference by causing a network to explore previously inaccessible states. Optimal perturbations break spin-spin correlations within a network, increasing the information available for inference and thus reducing sampling complexity by orders of magnitude. Our active learning framework could be powerful in the analysis of complex networks as well as in the rational design of experiments.
6M7WN4AX
webpage
Saved 2026-06-27
USG4B25T
computerProgram
A. V. Aditya
2026
Saved 2026-06-26
A plug-in debugger and visualizer for RL reward functions. Detects reward hacking, tracks training health, and renders a live terminal dashboard.
QKQ2D7C6
computerProgram
Prabhat Prakash
2026
Saved 2026-06-26
Tutorial to run basic DFT, MD and surface science calculations, in a classroom. Implementing minimal tools, lite exposure to pyscf, ORCA, Jaguar, lammps, Gromacs, Quantum Espresso and VASP codes
YGA4ZWFY
journalArticle
Adri C. T. van Duin, Siddharth Dasgupta, Francois Lorant, William A. Goddard
2001 · The Journal of Physical Chemistry A
Saved 2026-06-25
YW4WQ5KG
computerProgram
syt
2026
Saved 2026-06-25
Download PDF from Sci-Hub automatically For Zotero7
C4PTEUMI
journalArticle
Zhuoran Qiao, Matthew Welborn, Animashree Anandkumar, Frederick R. Manby, Thomas F. Miller
2020 · The Journal of Chemical Physics
Saved 2026-06-25
We introduce a machine learning method in which energy solutions from the Schrodinger equation are predicted using symmetry adapted atomic orbitals features and a graph neural-network architecture. \textsc{OrbNet} is shown to outperform existing methods in terms of learning efficiency and transferability for the prediction of density functional theory results while employing low-cost features that are obtained from semi-empirical electronic structure calculations. For applications to datasets of drug-like molecules, including QM7b-T, QM9, GDB-13-T, DrugBank, and the conformer benchmark dataset of Folmsbee and Hutchison, \textsc{OrbNet} predicts energies within chemical accuracy of DFT at a computational cost that is thousand-fold or more reduced.
MJ6Y69FS
forumPost
KatjaGrace
2026
Saved 2026-06-25
3JW9MN23
forumPost
dynomight
2026
Saved 2026-06-24
4NAKXPB2
preprint
Zhangde Song, Jieyu Lu, Yuanqi Du, Botao Yu, Thomas M. Pruyn, Yue Huang, Kehan Guo, Xiuzhe Luo et al.
2026
Saved 2026-06-24
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific "superintelligence". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.
56H7SUPA
preprint
Zhangde Song, Jieyu Lu, Yuanqi Du, Botao Yu, Thomas M. Pruyn, Yue Huang, Kehan Guo, Xiuzhe Luo et al.
2026
Saved 2026-06-24
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific "superintelligence". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.
JTC7P7A5
journalArticle
Yangrui Hu, Yi Hong Teoh, William Witczak-Krempa, Roger G Melko
2025 · Machine Learning: Science and Technology
Saved 2026-06-24
Abstract We explore the interplay of quantum computing and machine learning to advance experimental protocols for observing measurement-induced phase transitions (MIPTs) in quantum devices. In particular, we focus on trapped ion monitored circuits and apply the cross entropy benchmark recently introduced by Li et al (2023 Phys. Rev. Lett. 130 220404), which can mitigate the post-selection problem. By doing so, we reduce the number of projective measurements—the sample complexity—required per random circuit realization, which is a critical limiting resource in real devices. Since these projective measurement outcomes form a classical probability distribution, they are suitable for learning with a standard machine learning generative model. In this paper, we use a recurrent neural network to learn a representation of the measurement record for a native trapped-ion MIPT, and show that using this generative model can substantially reduce the number of measurements required to accurately estimate the cross entropy. This illustrates the potential of combining quantum computing and machine learning to overcome practical challenges in realizing quantum experiments.
J58LCN6R
webpage
Thinking Machines Lab
Saved 2026-06-24
On-policy, dense supervision is a useful tool for distillation
PIMKECYV
webpage
2009
Saved 2026-06-24
lukeprog's profile on LessWrong — A community blog devoted to refining the art of rationality
AWGUQYSS
forumPost
lukeprog
2011
Saved 2026-06-24
AJPT6PA4
journalArticle
Gwern
2016
Saved 2026-06-24
AIs limited to pure computation (Tool AIs) supporting humans, will be less intelligent, efficient, and economically valuable than more autonomous reinforcement-learning AIs (Agent AIs) who act on their own and meta-learn, because all problems are reinforcement-learning problems.
IW8CSLPB
webpage
Jiefang Xiao, Maolin Gao, Simon Weber, Guandao Yang, Daniel Cremers
2026
Saved 2026-06-24
Learning mappings between infinite-dimensional function spaces, or operator learning, is essential for many machine learning applications. Although transformer-based operators are popular, they often rely on token-wise attention. These methods treat continuous fields as discrete tokens and usually ignore the global functional structure. We introduce \emph{Functional Attention}, which reinterprets attention as a functional correspondence between adaptive bases. Inspired by geometric functional maps, our method replaces softmax affinities with structured linear operators. This yields a compact, generalizable, resolution-invariant representation that explicitly captures global dependencies. Experiments demonstrate that \emph{Functional Attention} can match state-of-the-art performance in many operator learning tasks, including solving PDEs, 3D segmentation, and regression, while remaining robust to varying discretizations. Project page is available at https://github.com/xjffff/FUNCATTN.
TS54NSXH
webpage
Saved 2026-06-24
RNKDCMCD
preprint
Terence Parr, Jeremy Howard
2018
Saved 2026-06-23
This paper is an attempt to explain all the matrix calculus you need in order to understand the training of deep neural networks. We assume no math knowledge beyond what you learned in calculus 1, and provide links to help you refresh the necessary math where needed. Note that you do not need to understand this material before you start learning to train and use deep learning in practice; rather, this material is for those who are already familiar with the basics of neural networks, and wish to deepen their understanding of the underlying math. Don't worry if you get stuck at some point along the way---just go back and reread the previous section, and try writing down and working through some examples. And if you're still stuck, we're happy to answer your questions in the Theory category at forums.fast.ai. Note: There is a reference section at the end of the paper summarizing all the key matrix calculus rules and terminology discussed here. See related articles at http://explained.ai
ZCB58J48
webpage
Saved 2026-06-23
A collaborative AI workspace, built on your company context. Build and orchestrate agents right alongside your team's projects, meetings, and connected apps.
ZTMITKGH
webpage
Saved 2026-06-23
A collaborative AI workspace, built on your company context. Build and orchestrate agents right alongside your team's projects, meetings, and connected apps.
HNKWXUU4
webpage
Saved 2026-06-23
RSK9F7ZW
forumPost
Philip Kiely [@philipkiely]
2026
Saved 2026-06-23
92NRTJVZ
webpage
Saved 2026-06-23
JYCZY28D
blogPost
Freddie deBoer
2023
Saved 2026-06-22
my descendants can get a job
V6QQFCXP
webpage
Saved 2026-06-22
MNPPIZFJ
journalArticle
Giacomo Torlai, Guglielmo Mazzola, Juan Carrasquilla, Matthias Troyer, Roger Melko, Giuseppe Carleo
2018 · Nature Physics
Saved 2026-06-22
The experimental realization of increasingly complex synthetic quantum systems calls for the development of general theoretical methods, to validate and fully exploit quantum resources. Quantum-state tomography (QST) aims at reconstructing the full quantum state from simple measurements, and therefore provides a key tool to obtain reliable analytics. Brute-force approaches to QST, however, demand resources growing exponentially with the number of constituents, making it unfeasible except for small systems. Here we show that machine learning techniques can be efficiently used for QST of highly-entangled states, in both one and two dimensions. Remarkably, the resulting approach allows one to reconstruct traditionally challenging many-body quantities - such as the entanglement entropy - from simple, experimentally accessible measurements. This approach can benefit existing and future generations of devices ranging from quantum computers to ultra-cold atom quantum simulators.
R62NCRNT
webpage
Saved 2026-06-21
JAXNTJIR
preprint
Ofir Press, Noah A. Smith, Mike Lewis
2022
Saved 2026-06-21
Since the introduction of the transformer model by Vaswani et al. (2017), a fundamental question has yet to be answered: how does a model achieve extrapolation at inference time for sequences that are longer than it saw during training? We first show that extrapolation can be enabled by simply changing the position representation method, though we find that current methods do not allow for efficient extrapolation. We therefore introduce a simpler and more efficient position method, Attention with Linear Biases (ALiBi). ALiBi does not add positional embeddings to word embeddings; instead, it biases query-key attention scores with a penalty that is proportional to their distance. We show that this method trains a 1.3 billion parameter model on input sequences of length 1024 that extrapolates to input sequences of length 2048, achieving the same perplexity as a sinusoidal position embedding model trained on inputs of length 2048 but training 11% faster and using 11% less memory. ALiBi's inductive bias towards recency also leads it to outperform multiple strong position methods on the WikiText-103 benchmark.
QJTRMLQ6
journalArticle
Hailan Ma, Zhenhong Sun, Daoyi Dong, Dong Gong
2026 · IEEE Transactions on Emerging Topics in Computational Intelligence
Saved 2026-06-21
Quantum state tomography (QST) is the process of reconstructing the complete state of a quantum system (mathematically described as a density matrix) through a series of different measurements. These measurements are performed on a number of identical copies of the quantum system, with outcomes gathered as probabilities/frequencies. QST aims to recover the density matrix and the corresponding properties of the quantum state from the measured frequencies. Although an informationally complete set of measurements can specify the quantum state accurately in an ideal scenario with a large number of identical copies, both the measurements and identical copies are restricted and imperfect in practical scenarios, making QST highly ill-posed. The conventional QST methods usually assume adequate or accurate measured frequencies or rely on manually designed regularizers to handle the ill-posed reconstruction problem, suffering from limited applications in realistic scenarios. Recent advances in deep neural networks (DNNs) led to the emergence of deep learning (DL) in QST. However, existing DL-based QST approaches often employ generic DNN models that are not optimized for imperfect conditions of QST. In this paper, we propose a transformerbased autoencoder architecture tailored for QST with imperfect measurement data. Our method leverages a transformer-based encoder to extract an informative latent representation (ILR) from imperfect measurement data and employs a decoder to predict the quantum states based on the ILR. We anticipate that the high-dimensional ILR will capture more comprehensive information about the quantum states. To achieve this, we conduct pre-training of the encoder using a pretext task that involves reconstructing high-quality frequencies from measured frequencies. Extensive simulations and experiments demonstrate the remarkable ability of the informative latent representation to deal with imperfect measurement data in QST.
PJ7M2ZUX
forumPost
Zvi
2023
Saved 2026-06-20
BJKMM4UB
blogPost
Zvi Mowshowitz
2022
Saved 2026-06-20
Or: Against Car Seat Laws At Least Beyond Age 2
QVB66AFX
webpage
Zilan Qian
2026
Saved 2026-06-19
and what it all means
ED6ES2UQ
webpage
Farhan Thawar
2026
Saved 2026-06-19
Introduce the Canadian AI Subscription Deduction (CASD) for individuals: a 100 percent personal tax deduction of up to $3,000 per year for qualifying AI subscriptions, learning tools, and productivity services.
7TWAXYTE
webpage
Saved 2026-06-18
A New Era of Midjourney, announcing Midjourney Medical: full-body Ultrasonic CT and the Midjourney Spa.
L6TCG85B
webpage
Saved 2026-06-18
Capture every rollout, score what happened, and retrain on the parts that matter.
4DEG8XK2
preprint
Amil Dravid, Yasaman Bahri, Alexei A. Efros, Yossi Gandelsman
2026
Saved 2026-06-18
We investigate whether neuron populations within neural networks evolve predictably with scale, extending scaling laws beyond macroscopic observables such as loss. To probe this question, we study Rosetta Neurons, a previously characterized class of neurons whose activation patterns are similar across independently trained models (Dravid et al., 2023). In separate analyses of language models up to 30B parameters and vision models up to 5B parameters, we observe that the population of Rosetta Neurons follows a sublinear power law in model size, growing in absolute number but occupying a shrinking fraction of the total neuron count. We further observe a Neuron Polarization Effect: Rosetta Neurons become more selective and increasingly monosemantic with scale, separating from a growing non-Rosetta population that remains less selective. An analytical model balancing feature utility against limited neuron capacity explains the sublinear power-law scaling and this polarization effect. Finally, we find that Rosetta Neurons become more domain-specialized with scale and illustrate their selectivity through a targeted data-filtering case study for continued pretraining. Our results point to a scaling law for interpretable, shared neuron-level structure, linking model size to systematic changes in neuron universality, selectivity, and specialization.
55J2BQMI
webpage
2026
Saved 2026-06-18
A Field Guide to AI Fellowships If you’re early in AI and you keep hearing “you should apply to a fellowship” without anyone telling you which one, for what, or how — this is for you.
NI2CCF23
forumPost
vivek [@itsreallyvivek]
2026
Saved 2026-06-18
JVXM69UA
forumPost
Elizabeth
2025
Saved 2026-06-17
Y6CDCR8G
journalArticle
Gwern
2026
Saved 2026-06-17
How much to edit? A top-𝑘 attention-window toy model of the publish-polish trade off: You should polish in inverse proportion to how much of your writing a reader will see; the more diverse your writing, the better off you are writing 𝑚𝑜𝑟𝑒 rather than 𝑏𝑒𝑡𝑡𝑒𝑟.
FTGG63SB
journalArticle
Monitoring Ultrafast Lattice Dynamics in 2D NbTe2
Christian Viernes
Saved 2026-06-17
QZE7IBUP
journalArticle
Eoin Whelan
2026 · Cyberpsychology, Behavior, and Social Networking
Saved 2026-06-17
The issue of whether social media use does or does not influence adolescent well-being remains a pressing concern for policymakers, parents, and researchers. As evidence of the harmful effects of social media, some point to the fact that young adults frequently regret the time they spent on social media when they were younger. However, highlighting a single instance of retrospective regret as a proxy for harm may be misleading. Our objective in this study is to provide a more accurate and contextually grounded assessment of social media regrets by benchmarking them against a broader set of common teenage regrets, as rated by young adults. We then determine which teenage regrets predict current life satisfaction. Four hundred young adults aged 20–24 from Ireland, the United Kingdom, and the United States were recruited via the Prolific platform and completed an online survey assessing the potency of 20 retrospective regrets when they reflect on their teenage years. Our findings indicate that social media was not a prominent source of regret when young adults reflect on their teenage years. In addition, regrets about social media use were not associated with current life satisfaction, but other regrets were. As such, our findings are consistent with other prior studies suggesting that at the population level, the harmful effects of social media use on adolescent well-being may be overstated. However, these results should be interpreted with caution and may not reflect the experience of all individuals within this population.
9BWCMH9H
preprint
Michael Broughton, Guillaume Verdon, Trevor McCourt, Antonio J. Martinez, Jae Hyeon Yoo, Sergei V. Isakov, Philip Massey, Ramin Halavati et al.
2021
Saved 2026-06-16
We introduce TensorFlow Quantum (TFQ), an open source library for the rapid prototyping of hybrid quantum-classical models for classical or quantum data. This framework offers high-level abstractions for the design and training of both discriminative and generative quantum models under TensorFlow and supports high-performance quantum circuit simulators. We provide an overview of the software architecture and building blocks through several examples and review the theory of hybrid quantum-classical neural networks. We illustrate TFQ functionalities via several basic applications including supervised learning for quantum classification, quantum control, simulating noisy quantum circuits, and quantum approximate optimization. Moreover, we demonstrate how one can apply TFQ to tackle advanced quantum learning tasks including meta-learning, layerwise learning, Hamiltonian learning, sampling thermal states, variational quantum eigensolvers, classification of quantum phase transitions, generative adversarial networks, and reinforcement learning. We hope this framework provides the necessary tools for the quantum computing and machine learning research communities to explore models of both natural and artificial quantum systems, and ultimately discover new quantum algorithms which could potentially yield a quantum advantage.
CQSE2IRX
webpage
2009
Saved 2026-06-16
H7UTFHSY
blogPost
Raghuveer Parthasarathy
2026
Saved 2026-06-16
What makes a college course popular or unpopular? I’ve long been interested in courses for non-science majors that satisfy “general education” requirements, their aim being to fos…
DBIIQDYN
webpage
Saved 2026-06-16
97LA4BML
preprint
Bruno Gavranović, Paul Lessard, Andrew Dudzik, Tamara von Glehn, João G. M. Araújo, Petar Veličković
2024
Saved 2026-06-16
We present our position on the elusive quest for a general-purpose framework for specifying and studying deep learning architectures. Our opinion is that the key attempts made so far lack a coherent bridge between specifying constraints which models must satisfy and specifying their implementations. Focusing on building a such a bridge, we propose to apply category theory -- precisely, the universal algebra of monads valued in a 2-category of parametric maps -- as a single theory elegantly subsuming both of these flavours of neural network design. To defend our position, we show how this theory recovers constraints induced by geometric deep learning, as well as implementations of many architectures drawn from the diverse landscape of neural networks, such as RNNs. We also illustrate how the theory naturally encodes many standard constructs in computer science and automata theory.
8R675IYK
blogPost
Joshua Gans
2026
Saved 2026-06-16
Straight out of Magic
V59CHHQ3
bookSection
Tanmay Sah, Vishal Srivastava, Dolly Sah, Kayden Jordan
2026 · Proceedings of the ACM Conference on AI and Agentic Systems · Association for Computing Machinery
Saved 2026-06-14
We study how runtime enforcement against unsafe actions affects end-to-end task performance in multi-step tool using large language model (LLM) agents. Using τ -bench across Airline and Retail domains, we compare baseline Tool-Calling, planning-integrated (Triad), and policy-mediated (Triad-Safety) architectures with GPT-OSS-20B and GLM-4-9B. We identify model dependent interaction horizons (15–30 turns) and decompose outcomes into overall success rate (SR), safe success rate (SSR), and unsafe success rate (USR). Our results reveal a persistent “Safety-Capability Gap”. While safety mediation can intercept up to 94% of non-compliant actions, it rarely translates into strictly safe goal attainment (SSR < 5% in most settings). We find that high unsafe success rates are primarily driven by “Integrity Leaks,” where models hallucinate user identifiers to bypass mandatory authentication. Recovery rates following blocked actions are consistently low, ranging from 21% for GPT-OSS-20B in simpler procedural tasks to near 0% in complex Retail scenarios. These results demonstrate that runtime enforcement imposes a significant “verifier tax” on conversational length and compute cost without guaranteeing safe completion, highlighting the critical need for agents capable of grounded identity verification and post-intervention reasoning.
97UXFNH4
blogPost
Tyler Cowen
2022
Saved 2026-06-14
Teenage mental health has been a source of growing concern over the past decade, with recent whistleblower testimony pointing to the mental health risks of spending time on social media platforms, especially for girls. This paper investigates the extent to which social media are harmful for teenagers, leveraging rich administrative data from the Canadian province […]
WE59LKIY
preprint
Yarin Gal, Zoubin Ghahramani
2016
Saved 2026-06-13
Deep learning tools have gained tremendous attention in applied machine learning. However such tools for regression and classification do not capture model uncertainty. In comparison, Bayesian models offer a mathematically grounded framework to reason about model uncertainty, but usually come with a prohibitive computational cost. In this paper we develop a new theoretical framework casting dropout training in deep neural networks (NNs) as approximate Bayesian inference in deep Gaussian processes. A direct result of this theory gives us tools to model uncertainty with dropout NNs -- extracting information from existing models that has been thrown away so far. This mitigates the problem of representing uncertainty in deep learning without sacrificing either computational complexity or test accuracy. We perform an extensive study of the properties of dropout's uncertainty. Various network architectures and non-linearities are assessed on tasks of regression and classification, using MNIST as an example. We show a considerable improvement in predictive log-likelihood and RMSE compared to existing state-of-the-art methods, and finish by using dropout's uncertainty in deep reinforcement learning.
YQ3F63JJ
webpage
Saved 2026-06-12
What happens in a world where AIs make scientific discoveries that humans cannot understand?
G8YUKWHH
conferencePaper
Susumu Kuno, Anthony G. Oettinger
1963 · Proceedings of the November 12-14, 1963, fall joint computer conference on XX - AFIPS '63 (Fall) · ACM Press
Saved 2026-06-12
Z5492T9W
webpage
Saved 2026-06-12
YTLL8HPM
webpage
Daan Juijn, Stan van Baarsen, Judith Dada, Lily Stelling, Philip Fox, Alex Petropoulos, Michiel Bakker
Saved 2026-06-12
A five-year scenario about AI and Europe's impending slide into irrelevance, with a 2034 epilogue that describes how the collapse of the European model could have been prevented.
JFTBBIGW
webpage
Saved 2026-06-12
Learn faster, achieve more
3E3URGC8
webpage
Saved 2026-06-12
Jeri Ellsworth documents her amateur science experiments.
6D9S2D29
blogPost
2018
Saved 2026-06-12
FUMMYSYE
book
Functions 11
Saved 2026-06-11
RFXY7AYC
webpage
2026
Saved 2026-06-11
By Ryan Lopopolo, Member of the Technical Staff
U7NY6RHF
webpage
2026
Saved 2026-06-11
A case against measuring AI-assisted engineering by code volume and other vanity metrics instead of outcomes.
VPLW9QN9
journalArticle
How Does Batch Normalization Help Optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, Aleksander Ma
Saved 2026-06-11
Batch Normalization (BatchNorm) is a widely adopted technique that enables faster and more stable training of deep neural networks (DNNs). Despite its pervasiveness, the exact reasons for BatchNorm’s effectiveness are still poorly understood. The popular belief is that this effectiveness stems from controlling the change of the layers’ input distributions during training to reduce the so-called “internal covariate shift”. In this work, we demonstrate that such distributional stability of layer inputs has little to do with the success of BatchNorm. Instead, we uncover a more fundamental impact of BatchNorm on the training process: it makes the optimization landscape significantly smoother. This smoothness induces a more predictive and stable behavior of the gradients, allowing for faster training.
5KL6GC7U
forumPost
luuk0987
2018
Saved 2026-06-11
MMQCV9WV
webpage
Saved 2026-06-10
Repository for the book "Crafting Interpreters". Contribute to munificent/craftinginterpreters development by creating an account on GitHub.
XAZAR9WW
webpage
Saved 2026-06-10
3VWBJEJ8
webpage
Saved 2026-06-10
M7B78Y4F
journalArticle
NBER WORKING PAPER SERIES
Caitlin K Myers, Ezekiel Hooper
Saved 2026-06-09
The U.S. general fertility rate has fallen by 22% since 2007, a sustained decline not readily explained by economic conditions, contraceptive use, housing or childcare costs, or other commonly cited factors. We assess the potential role of a different shock: the diffusion of the smartphone. The U.S. rollout of the iPhone, the first modern smartphone, provides a natural experiment: from June 2007 through February 2011, the device was sold only on AT&T, allowing us to identify its effect from variation in AT&T’s mobile broadband coverage. Entropy-balanced Poisson and synthetic difference-in-differences event studies imply that access to the iPhone reduced births by 4.5–8.0% at ages 15–19 and 3.2–6.6%at ages 20–24, with statistically significant but smaller declines among older cohorts. Placebo analyses applied to Verizon and Sprint’s pre-2011 coverage footprint are null. Taken together, these cohort effects imply that the diffusion of the iPhone deepened the decline in births among women under 30 while suppressing the rise in births among older women. Overall, the diffusion of the iPhone explains 33–52% of the decline in the general fertility rate among women aged 15–44. National-survey evidence on time use and sexual behavior is consistent with the iPhone reducing in-person interactions, increasing pornography use, and reducing sexual frequency.
HC6KGQNV
blogPost
Tyler Cowen
2026
Saved 2026-06-09
The U.S. general fertility rate has fallen by 22% since 2007, a sustained decline not readily explained by economic conditions, contraceptive use, housing or childcare costs, or other commonly cited factors. We assess the potential role of a different shock: the diffusion of the smartphone. The U.S. rollout of the iPhone, the first modern smartphone, […]
GZE4BQ5V
webpage
german s
2026
Saved 2026-06-08
The act of pumping immense, disproportionate resources into a previously casual or complex and layered activity to forcefully extract and squeeze out the purest, most concentrated dopamine hit
E6JBKC69
webpage
Saved 2026-06-08
Five exercises for building research taste (and three failure modes).
54GF6Z42
webpage
Saved 2026-06-07
5KPHIB45
journalArticle
Clement Gosselin, Thierry Laliberte, Audrey Veillette
2015 · IEEE Transactions on Robotics
Saved 2026-06-06
VZCHL6NC
webpage
Saved 2026-06-05
Our progress toward recursive self-improvement, and its implications.
VJSDI3PR
preprint
Ali Kayyam, Anusha Madan Gopal, M Anthony Lewis
2026
Saved 2026-06-05
Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a central role. However, the individual contribution of these three projections and the impact of omitting some remain poorly understood. We systematically evaluate three projection sharing constraints: a) Q-K=V (shared key-value), b) Q=K-V (shared query-key), and c) Q=K=V (single projection). The last two variants produce symmetric attention maps; to address this, we also explore asymmetric attention via 2D positional encodings. Through experiments spanning synthetic tasks, vision (MNIST, CIFAR, TinyImageNet, anomaly), and language modeling (300M and 1.2B parameter models on 10B tokens), we discovered that our transformers perform on par or occasionally better than the QKV transformer. In language modeling, Q-K=V projection sharing achieves 50% KV cache reduction with only 3.1% perplexity degradation. Crucially, projection sharing is complementary to head sharing (GQA/MQA): combining Q-K=V with GQA-4 yields 87.5% cache reduction, while Q-K=V + MQA achieves 96.9%, enabling practical on-device inference. We show that Q-K=V preserves quality because keys and values can occupy similar representational spaces and attention operates in a low-rank regime, whereas Q=K-V breaks attention directionality. Our results systematically characterize projection sharing as an underexplored instance of weight tying in attention, with direct, quantifiable inference memory benefits, particularly valuable for edge deployment. The code is publicly available at https://github.com/Brainchip-Inc/Do-Transformers-Need-3-Projections
TKV26IVD
forumPost
will brown [@willccbb]
2026
Saved 2026-06-05
LS6HK7DE
presentation
Post-Training Distillation for LLMs
Rishabh Agarwal
Saved 2026-06-05
2DAWV6WY
journalArticle
Felipe L. Gewers, Gustavo R. Ferreira, Henrique F. de Arruda, Filipi N. Silva, Cesar H. Comin, Diego R. Amancio, Luciano da F. Costa
2022 · ACM Computing Surveys
Saved 2026-06-04
Principal component analysis (PCA) is often used for analyzing data in the most diverse areas. In this work, we report an integrated approach to several theoretical and practical aspects of PCA. We start by providing, in an intuitive and accessible manner, the basic principles underlying PCA and its applications. Next, we present a systematic, though no exclusive, survey of some representative works illustrating the potential of PCA applications to a wide range of areas. An experimental investigation of the ability of PCA for variance explanation and dimensionality reduction is also developed, which confirms the efficacy of PCA and also shows that standardizing or not the original data can have important effects on the obtained results. Overall, we believe the several covered issues can assist researchers from the most diverse areas in using and interpreting PCA.
5FJ8GCK8
webpage
daniel.graham@ourmedia.co.uk
2026
Saved 2026-06-04
Footage captured in Australia shows one of the world's most peculiar animals playing in a creek.
K677FZMC
webpage
Substack
Saved 2026-06-04
QJUG8HUU
webpage
Tomas Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean
2013
Saved 2026-06-03
We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuring syntactic and semantic word similarities.
Y2ICP9WA
webpage
Jake Tae
2020
Saved 2026-06-03
In a previous post, we discussed how we can use tf-idf vectorization to encode documents into vectors. While probing more into this topic and geting a taste of what NLP is like, I decided to take a jab at another closely related, classic topic in NLP: word2vec. word2vec is a technique introduced by Google engineers in 2013, popularized by statements such as “king - man + woman = queen.” The gist of it, as you may know, is that we can express words as vectors that encode their semantics in a meaningful way.
44IWHFQJ
webpage
Saved 2026-06-03
86546CRL
webpage
2026
Saved 2026-06-03
The generational collapse in literacy is measurable, persistent, and likely to get worse.
YV7NGZRP
webpage
Saved 2026-06-02
I turned 30 last week and a friend asked me if I'd figured out any life advice in the past decade worth passing on.  I'm somewhat hesitant to publish this because I think these lists usually seem...
NMK5NMNM
forumPost
Eliezer Yudkowsky
2018
Saved 2026-06-02
UCDJEELJ
webpage
Hear This Idea
Saved 2026-06-02
Hear This Idea
AHI5ZBYF
bookSection
Introduction to renormalization group methods in physics / R.J. Creswick, H.A. Farach, C.P. Poole
Richard J. Creswick, Horacio A. Farach, Charles P. Poole
1992 · Introduction to renormalization group methods in physics · Wiley
Saved 2026-06-02
56ZYHPY6
forumPost
Ruby
2023
Saved 2026-06-02
87Y5ZFWQ
forumPost
Raemon, RobertM, habryka
2019
Saved 2026-06-02
C9K2AXIF
forumPost
Eliezer Yudkowsky
2018
Saved 2026-06-02
EIREVGQT
forumPost
atlas [@creatine_cycle]
2026
Saved 2026-06-02
X6X4KRY5
journalArticle
Stephen Wolfram
2023 · Stephen Wolfram Writings
Saved 2026-06-02
Follow Stephen Wolfram's evolution of research on the 2nd Law from early interest as a 12-year-old to his current findings in the Wolfram Physics Project. View a large collection of images of early and published works.
75YJDE92
forumPost
Simon Berens
2023
Saved 2026-06-02
WGFY3RAR
conferencePaper
Chris Bamford, Simon M. Lucas
2020 · 2020 IEEE Conference on Games (CoG) · IEEE
Saved 2026-06-01
WY8A63IS
webpage
Saved 2026-06-01
RHEAMIC5
preprint
Thaddäus Wiedemer, Yuxuan Li, Paul Vicol, Shixiang Shane Gu, Nick Matarese, Kevin Swersky, Been Kim, Priyank Jaini et al.
2025
Saved 2026-06-01
The remarkable zero-shot capabilities of Large Language Models (LLMs) have propelled natural language processing from task-specific models to unified, generalist foundation models. This transformation emerged from simple primitives: large, generative models trained on web-scale data. Curiously, the same primitives apply to today's generative video models. Could video models be on a trajectory towards general-purpose vision understanding, much like LLMs developed general-purpose language understanding? We demonstrate that Veo 3 can solve a broad variety of tasks it wasn't explicitly trained for: segmenting objects, detecting edges, editing images, understanding physical properties, recognizing object affordances, simulating tool use, and more. These abilities to perceive, model, and manipulate the visual world enable early forms of visual reasoning like maze and symmetry solving. Veo's emergent zero-shot capabilities indicate that video models are on a path to becoming unified, generalist vision foundation models.
WK92AV4J
preprint
Aaron van den Oord, Oriol Vinyals, Koray Kavukcuoglu
2018
Saved 2026-06-01
Learning useful representations without supervision remains a key challenge in machine learning. In this paper, we propose a simple yet powerful generative model that learns such discrete representations. Our model, the Vector Quantised-Variational AutoEncoder (VQ-VAE), differs from VAEs in two key ways: the encoder network outputs discrete, rather than continuous, codes; and the prior is learnt rather than static. In order to learn a discrete latent representation, we incorporate ideas from vector quantisation (VQ). Using the VQ method allows the model to circumvent issues of "posterior collapse" -- where the latents are ignored when they are paired with a powerful autoregressive decoder -- typically observed in the VAE framework. Pairing these representations with an autoregressive prior, the model can generate high quality images, videos, and speech as well as doing high quality speaker conversion and unsupervised learning of phonemes, providing further evidence of the utility of the learnt representations.
G6CJWLWJ
computerProgram
Jonathan Heek, Anselm Levskaya, Avital Oliver, Marvin Ritter, Bertrand Rondepierre, Andreas Steiner, Marc van Zee
2024
Saved 2026-06-01
NV7QZTSB
computerProgram
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Yash Katariya, Chris Leary, Dougal Maclaurin, George Necula et al.
2018
Saved 2026-06-01
KJ4AFF4X
webpage
Thinking Machines Lab
Saved 2026-06-01
Proposal requirements, selection criteria, timeline, and terms for the Interactivity Research Grant Program.
WE6F5NCV
preprint
Ruta Binkyte, Ivaxi Sheth, Zhijing Jin, Mohammad Havaei, Bernhard Schölkopf, Mario Fritz
2026
Saved 2026-05-31
As artificial intelligence (AI), including machine learning (ML) models and foundation models (FMs), is increasingly deployed in high-stakes domains, ensuring their trustworthiness has become a central challenge. However, the core trustworthy AI objectives, such as fairness, robustness, privacy, and explainability, are hard to achieve simultaneously, especially while preserving utility. This position paper argues that causality is necessary to understand and balance trade-offs in performance and multiple objectives of trustworthy AI. We ground our arguments in re-interpreting trustworthy AI trade-offs as incompatible invariance requirements under different changes to the data-generating process. We then illustrate that causality provides a unifying framework for understanding how trade-offs in trustworthy AI arise, and how they can be softened or resolved through selective invariance. This perspective applies to both classical ML models and large-scale FMs. Our paper discusses how causal assumptions may be applied explicitly or implicitly in modern large-scale systems. Finally, we outline open challenges and opportunities for using causality to build more trustworthy AI.
YZKY8H7I
preprint
Xuanqiang Angelo Huang, Charlie Tharas, Samuele Marro, Van Q. Truong, Bernhard Schölkopf, Emanuele La Malfa, Zhijing Jin
2026
Saved 2026-05-31
Ensuring that AI agents behave safely and beneficially when interacting with other parties has emerged as one of the central challenges of modern AI safety. While mechanism design, as the theory of designing rules to align individual and collective objectives, can incentivize cooperative behavior, it is still an open question whether it alone is sufficient to maximize LLM agents' social welfare. This work proves that the answer is negative: drawing from incomplete contract theory, we formally show that when contracts cannot distinguish all relevant future contingencies, there is a strictly positive welfare loss that no realistic mechanism can eliminate. We show that prosocial agents, who weigh others' welfare alongside their own, can close this gap and achieve outcomes that are socially superior and individually beneficial. Experimentally, we show that in multi-agent resource-allocation environments and canonical social dilemmas where agents are powered by large language models, prosociality is beneficial. The implication for AI safety is clear: to enable cooperative interactions at scale, designing adequate mechanisms is not sufficient; agents must be built to be intrinsically prosocial.
WAXYXXMQ
preprint
Jiarui Liu, Lechen Zhang, Yongjin Yang, Yinghui He, Yingheng Wang, Weihao Xuan, Zhijing Jin, Mona Diab
2026
Saved 2026-05-31
Supervised fine-tuning (SFT) is widely used to inject new knowledge into language models, but it often degrades pretrained capabilities such as reasoning and general-domain performance. We argue this forgetting arises because fine-tuning targets from humans or external systems diverge from the model's autoregressive distribution, forcing the optimizer to imitate low-probability token sequences. To address this problem, we propose MixSD, a simple external-teacher-free method for distribution-aligned knowledge injection. Instead of training on fixed targets, MixSD constructs supervision dynamically by mixing tokens from two conditionals of the base model itself: an expert conditional that observes the injected fact in context, and a naive conditional that reflects the model's original prior. The resulting supervision sequences preserve the factual learning signal while remaining substantially closer to the base model's distribution. We evaluate MixSD on two synthetic corpora that we construct to study factual recall and arithmetic function acquisition in a controlled setting, together with established benchmarks for open-domain factual question answering and knowledge editing. Across multiple model scales and settings, MixSD consistently achieves a better memorization-retention trade-off compared to SFT and on-policy self distillation baselines, retaining up to 100% of the base model's held-out capability while maintaining near-perfect training accuracy, whereas standard SFT retains as little as 1%. We further show that MixSD produces substantially lower-NLL supervision targets under the base model and reduces harmful movement along Fisher-sensitive parameter directions. These results suggest that aligning supervision with the model's native generation distribution is a simple and effective principle for knowledge injection that mitigates catastrophic forgetting.
M9ACEB7F
conferencePaper
Alexander Viand, Patrick Jattke, Anwar Hithnawi
2021 · 2021 IEEE Symposium on Security and Privacy (SP)
Saved 2026-05-31
Fully Homomorphic Encryption (FHE) allows a third party to perform arbitrary computations on encrypted data, learning neither the inputs nor the computation results. Hence, it provides resilience in situations where computations are carried out by an untrusted or potentially compromised party. This powerful concept was first conceived by Rivest et al. in the 1970s. However, it remained unrealized until Craig Gentry presented the first feasible FHE scheme in 2009.The advent of the massive collection of sensitive data in cloud services, coupled with a plague of data breaches, moved highly regulated businesses to increasingly demand confidential and secure computing solutions. This demand, in turn, has led to a recent surge in the development of FHE tools. To understand the landscape of recent FHE tool developments, we conduct an extensive survey and experimental evaluation to explore the current state of the art and identify areas for future development.In this paper, we survey, evaluate, and systematize FHE tools and compilers. We perform experiments to evaluate these tools’ performance and usability aspects on a variety of applications. We conclude with recommendations for developers intending to develop FHE-based applications and a discussion on future directions for FHE tools development.
2UHAC53G
conferencePaper
Hossein Shafagh, Lukas Burkhalter, Anwar Hithnawi, Simon Duquennoy
2017 · Proceedings of the 2017 on Cloud Computing Security Workshop · ACM
Saved 2026-05-31
G6GWEFGK
preprint
Hidde Lycklama, Lukas Burkhalter, Alexander Viand, Nicolas Küchler, Anwar Hithnawi
2023
Saved 2026-05-31
Even though recent years have seen many attacks exposing severe vulnerabilities in Federated Learning (FL), a holistic understanding of what enables these attacks and how they can be mitigated effectively is still lacking. In this work, we demystify the inner workings of existing (targeted) attacks. We provide new insights into why these attacks are possible and why a definitive solution to FL robustness is challenging. We show that the need for ML algorithms to memorize tail data has significant implications for FL integrity. This phenomenon has largely been studied in the context of privacy; our analysis sheds light on its implications for ML integrity. We show that certain classes of severe attacks can be mitigated effectively by enforcing constraints such as norm bounds on clients' updates. We investigate how to efficiently incorporate these constraints into secure FL protocols in the single-server setting. Based on this, we propose RoFL, a new secure FL system that extends secure aggregation with privacy-preserving input validation. Specifically, RoFL can enforce constraints such as $L_2$ and $L_\infty$ bounds on high-dimensional encrypted model updates.
F7HB8AZX
journalArticle
Tahsin Reza, Aaron Zimmer, Jose Manuel Delgado Blasco, Parwant Ghuman, Tanuj Kr Aasawat, Matei Ripeanu
2018 · IEEE Transactions on Parallel and Distributed Systems
Saved 2026-05-31
PLE3DZAR
preprint
Tahsin Reza, Hassan Halawa, Matei Ripeanu, Geoffrey Sanders, Roger Pearce
2020
Saved 2026-05-31
Pattern matching is a fundamental tool for answering complex graph queries. Unfortunately, existing solutions have limited capabilities: they do not scale to process large graphs and/or support only a restricted set of search templates or usage scenarios. We present an algorithmic pipeline that bases pattern matching on constraint checking. The key intuition is that each vertex or edge participating in a match has to meet a set of constrains implicitly specified by the search template. The pipeline we propose, generates these constraints and iterates over them to eliminate all the vertices and edges that do not participate in any match, and reduces the background graph to a subgraph which is the union of all matches. Additional analysis can be performed on this annotated, reduced graph, such as full match enumeration. Furthermore, a vertex-centric formulation for constraint checking algorithms exists, and this makes it possible to harness existing high-performance, vertex-centric graph processing frameworks. The key contribution of this work is a design following the constraint checking approach for exact matching and its experimental evaluation. We show that the proposed technique: (i) enables highly scalable pattern matching in labeled graphs, (ii) supports arbitrary patterns with 100% precision, (iii) always selects all vertices and edges that participate in matches, thus offering 100% recall, and (iv) supports a set of popular data analysis scenarios. We implement our approach on top of HavoqGT, an open-source asynchronous graph processing framework, and demonstrate its advantages through strong and weak scaling experiments on massive scale real-world (up to 257 billion edges) and synthetic (up to 4.4 trillion edges) labeled graphs respectively, and at scales (1,024 nodes / 36,864 cores), orders of magnitude larger than used in the past for similar problems.
9F6XQFMX
webpage
Saved 2026-05-31
3MCDZDM7
preprint
Benjamin Eyre, Elliot Creager, David Madras, Vardan Papyan, Richard Zemel
2023
Saved 2026-05-31
Designing deep neural network classifiers that perform robustly on distributions differing from the available training data is an active area of machine learning research. However, out-of-distribution generalization for regression-the analogous problem for modeling continuous targets-remains relatively unexplored. To tackle this problem, we return to first principles and analyze how the closed-form solution for Ordinary Least Squares (OLS) regression is sensitive to covariate shift. We characterize the out-of-distribution risk of the OLS model in terms of the eigenspectrum decomposition of the source and target data. We then use this insight to propose a method for adapting the weights of the last layer of a pre-trained neural regression model to perform better on input data originating from a different distribution. We demonstrate how this lightweight spectral adaptation procedure can improve out-of-distribution performance for synthetic and real-world datasets.
DIIAWRAF
preprint
Rushabh Solanki, Elliot Creager
2024
Saved 2026-05-31
The deployment of AI in consumer products is currently focused on the use of so-called foundation models, large neural networks pre-trained on massive corpora of digital records. This emphasis on scaling up datasets and pre-training computation raises the risk of further consolidating the industry, and enabling monopolistic (or oligopolistic) behavior. Judges and regulators seeking to improve market competition may employ various remedies. This paper explores dissolution -- the breaking up of a monopolistic entity into smaller firms -- as one such remedy, focusing in particular on the technical challenges and opportunities involved in the breaking up of large models and datasets. We show how the framework of Conscious Data Contribution can enable user autonomy during under dissolution. Through a simulation study, we explore how fine-tuning and the phenomenon of "catastrophic forgetting" could actually prove beneficial as a type of machine unlearning that allows users to specify which data they want used for what purposes.
HJTTMCP4
preprint
Rushabh Solanki, Meghana Bhange, Ulrich Aïvodji, Elliot Creager
2026
Saved 2026-05-31
The integration of AI into daily life has generated considerable attention and excitement, while also raising concerns about automating algorithmic harms and re-entrenching existing social inequities. While the responsible deployment of trustworthy AI systems is a worthy goal, there are many possible ways to realize it, from policy and regulation to improved algorithm design and evaluation. In fact, since AI trains on social data, there is even a possibility for everyday users, citizens, or workers to directly steer the AI system's behavior through Algorithmic Collective Action, by deliberately modifying the data they share with a platform to drive its learning process in their favor. This paper considers how these grassroots efforts to influence AI interact with methods used by AI firms and governments to improve model trustworthiness. In particular, we focus on the setting where the AI firm deploys a differentially private model, motivated by the growing regulatory focus on privacy and data protection. We investigate how the use of Differentially Private Stochastic Gradient Descent (DP-SGD) affects the collective's ability to influence the learning process. Our findings show that while differential privacy protects individual data, it introduces challenges for effective algorithmic collective action. We establish this trade-off formally by characterizing lower bounds on the success of algorithmic collective action under differential privacy as a function of the collective's size and the firm's privacy parameters. We then verify these trends experimentally by simulating collective action during the training of deep neural network classifiers across several datasets. Finally, we perform a stylized economic analysis of privacy costs to integrate additional incentives, analyzing how utility and participation costs influence the formation of collectives under private training regimes.
YGP6KSJL
journalArticle
Conscious Data Contribution via Community-Driven Chain-of-Thought Distillation
Lena Libon, Meghana Bhange, Rushabh Solanki, Elliot Creager, Ulrich Aïvodji
Saved 2026-05-31
The current era of AI development places a heavy emphasis on training large models on increasingly scaled-up datasets. This paradigm has catalyzed entirely new product categories, such as LLM chatbots, while also raising concerns about data privacy and consumer choice. In this paper, we consider questions of data portability and user autonomy in the context of LLMs that “reason” using chain-ofthought (CoT) traces, computing intermediate text artifacts from user input before producing a final output. We first interpret recent data privacy and portability law to argue that these intermediate computations qualify as users’ personal data. Then, building on the existing framework of Conscious Data Contribution, we show how communities who receive low utility from an available model can aggregate and distill their shared knowledge into an alternate model better aligned with their goals. We verify this approach empirically and investigate the effects of community diversity, reasoning granularity, and community size on distillation performance.
ASVTD49F
webpage
About
Saved 2026-05-31
5ZC3ACBS
preprint
Jonathan Gratus
2017
Saved 2026-05-31
In this article we present pictorially the foundation of differential geometry which is a crucial tool for multiple areas of physics, notably general and special relativity, but also mechanics, thermodynamics and solving differential equations. As all the concepts are presented as pictures, there are no equations in this article. As such this article may be read by pre-university students who enjoy physics, mathematics and geometry. However it will also greatly aid the intuition of an undergraduate and masters students, learning general relativity and similar courses. It concentrates on the tools needed to understand Maxwell's equations thus leading to the goal of presenting Maxwell's equations as 3 pictures.
VISCSM3S
journalArticle
SEMANTIC UNCERTAINTY: LINGUISTIC INVARIANCES FOR UNCERTAINTY ESTIMATION IN NATURAL LANGUAGE GENERATION
Lorenz Kuhn, Yarin Gal, Sebastian Farquhar
2023
Saved 2026-05-31
We introduce a method to measure uncertainty in large language models. For tasks like question answering, it is essential to know when we can trust the natural language outputs of foundation models. We show that measuring uncertainty in natural language is challenging because of ‘semantic equivalence’—different sentences can mean the same thing. To overcome these challenges we introduce semantic entropy—an entropy which incorporates linguistic invariances created by shared meanings. Our method is unsupervised, uses only a single model, and requires no modifications to ‘off-the-shelf’ language models. In comprehensive ablation studies we show that the semantic entropy is more predictive of model accuracy on question answering data sets than comparable baselines.
YDG3UAFP
blogPost
Steve Yegge
2026
Saved 2026-05-30
Today we will pour one out for the vaunted technical interview process, which is on its last leg. And we’ll talk a little about what’s…
T4DIRQX9
blogPost
Saved 2026-05-30
SDUHW7SK
blogPost
Terence Eden
2026
Saved 2026-05-29
It can be hard running a small business. If you want to sell to a large organisation like the UK Government, there are forms to fill in, checks to comply with, tenders to bid on, and a hundred other things. Luckily, there's the RM6237 Low Value Purchase System to make everything better. If a department wants to buy something below a certain threshold, they can contact any of the registered…
7AZVSBCU
webpage
Saved 2026-05-29
AI is doing to programming what framework-brain did to the frontend before. Deskilling, or just working at a higher level of abstraction?
ISZ5KHAK
webpage
Saved 2026-05-29
S4R8RV99
webpage
Saved 2026-05-29
DL7U2LWD
webpage
2026
Saved 2026-05-29
We implemented the entire LLM decode pass in a single persistent kernel, no kernel launches, no interruptions, achieving 3,000+ tokens/s per request on AMD MI300X.
JEQ5W3M2
webpage
2026
Saved 2026-05-29
Today, Kog AI launches a tech preview of the Kog Inference Engine (KIE): 3,000 output tokens/s per request on 8× AMD MI300X GPUs and 2,100 on 8× NVIDIA H200 (FP16, no speculative decoding). This preview runs a 2B model, with support for large third-party MoE models coming next at similar speeds.
CFHBESJF
webpage
Lorraine Boissoneault
Saved 2026-05-29
A new movie sets its doomed entrepreneurs amidst 17th-century “tulipmania”—but historians of the phenomenon have their own bubble to burst
L8QUMZBX
videoRecording
Mike Morrison
2019
Saved 2026-05-29
Creative Commons Attribution license (reuse allowed)
W8D9T3MV
blogPost
Zen Faulkes
2016
Saved 2026-05-29
7ZGEGTM5
blogPost
Saved 2026-05-28
CQ8RDG7U
webpage
2026
Saved 2026-05-28
Hundreds of University of California professors are urging the system to reinstate the SAT or ACT requirement for STEM majors by 2027, saying that the test-optional policy has created a widening preparation gap that threatens the value of UC science, technology, engineering and math degrees.
JYALY593
preprint
Daniel Uzcategui-Contreras, Antonio Guerra, Sebastian Niklitschek, Aldo Delgado
2024
Saved 2026-05-28
In this work, we propose a machine learning-based approach to address a specific aspect of the Quantum Marginal Problem: reconstructing a global density matrix compatible with a given set of quantum marginals. Our method integrates a quantum marginal imposition technique with convolutional denoising autoencoders. The loss function is carefully designed to enforce essential physical constraints, including Hermiticity, positivity, and normalization. Through extensive numerical simulations, we demonstrate the effectiveness of our approach, achieving high success rates and accuracy. Furthermore, we show that, in many cases, our model offers a faster alternative to state-of-the-art semidefinite programming solvers without compromising solution quality. These results highlight the potential of machine learning techniques for solving complex problems in quantum mechanics.
ABZL3RHB
journalArticle
Giacomo Torlai, Roger G. Melko
2017 · Physical Review Letters
Saved 2026-05-28
We present an algorithm for error correction in topological codes that exploits modern machine learning techniques. Our decoder is constructed from a stochastic neural network called a Boltzmann machine, of the type extensively used in deep learning. We provide a general prescription for the training of the network and a decoding strategy that is applicable to a wide variety of stabilizer codes with very little specialization. We demonstrate the neural decoder numerically on the well-known two dimensional toric code with phase-flip errors.
LE397YFH
journalArticle
Dmytro Bondarenko, Polina Feldmann
2020 · Physical Review Letters
Saved 2026-05-28
Entangled states are an important resource for quantum computation, communication, metrology, and the simulation of many-body systems. However, noise limits the experimental preparation of such states. Classical data can be efficiently denoised by autoencoders---neural networks trained in unsupervised manner. We develop a novel quantum autoencoder that successfully denoises Greenberger-Horne-Zeilinger states subject to spin-flip errors and random unitary noise. Various emergent quantum technologies could benefit from the proposed unsupervised quantum neural networks.
JKYTGZ5P
journalArticle
Adriano Macarone Palmieri, Guillem Müller-Rigat, Anubhav Kumar Srivastava, Maciej Lewenstein, Grzegorz Rajchel-Mieldzioć, Marcin Płodzień
2024 · Physical Review Research
Saved 2026-05-28
Resource-efficient quantum state tomography is one of the key ingredients of future quantum technologies. In this work, we propose a new tomography protocol combining standard quantum state reconstruction methods with an attention-based neural network architecture. We show how the proposed protocol is able to improve the averaged fidelity reconstruction over linear inversion and maximum-likelihood estimation in the finite-statistics regime, reducing at least by an order of magnitude the amount of necessary training data. We demonstrate the potential use of our protocol in physically relevant scenarios, in particular, to certify metrological resources in the form of many-body entanglement generated during the spin squeezing protocols. This could be implemented with the current quantum simulator platforms, such as trapped ions, and ultra-cold atoms in optical lattices.
CDQFQXY3
journalArticle
Shahnawaz Ahmed, Carlos Sánchez Muñoz, Franco Nori, Anton Frisk Kockum
2021 · Physical Review Letters
Saved 2026-05-28
Quantum state tomography (QST) is a challenging task in intermediate-scale quantum devices. Here, we apply conditional generative adversarial networks (CGANs) to QST. In the CGAN framework, two duelling neural networks, a generator and a discriminator, learn multi-modal models from data. We augment a CGAN with custom neural-network layers that enable conversion of output from any standard neural network into a physical density matrix. To reconstruct the density matrix, the generator and discriminator networks train each other on data using standard gradient-based methods. We demonstrate that our QST-CGAN reconstructs optical quantum states with high fidelity orders of magnitude faster, and from less data, than a standard maximum-likelihood method. We also show that the QST-CGAN can reconstruct a quantum state in a single evaluation of the generator network if it has been pre-trained on similar quantum states.
K7RSADJK
journalArticle
Adriano Macarone Palmieri, Guillem Müller-Rigat, Anubhav Kumar Srivastava, Maciej Lewenstein, Grzegorz Rajchel-Mieldzioć, Marcin Płodzień
2024 · Physical Review Research
Saved 2026-05-28
Resource-efficient quantum state tomography is one of the key ingredients of future quantum technologies. In this work, we propose a new tomography protocol combining standard quantum state reconstruction methods with an attention-based neural network architecture. We show how the proposed protocol is able to improve the averaged fidelity reconstruction over linear inversion and maximum-likelihood estimation in the finite-statistics regime, reducing at least by an order of magnitude the amount of necessary training data. We demonstrate the potential use of our protocol in physically relevant scenarios, in particular, to certify metrological resources in the form of many-body entanglement generated during the spin squeezing protocols. This could be implemented with the current quantum simulator platforms, such as trapped ions, and ultra-cold atoms in optical lattices.
HC4Y8VMX
journalArticle
Sanjaya Lohani, Brian T. Kirby, Michael Brodsky, Onur Danaci, Ryan T. Glasser
2020 · Machine Learning: Science and Technology
Saved 2026-05-27
We build a general quantum state tomography framework that makes use of machine learning techniques to reconstruct quantum states from a given set of coincidence measurements. For a wide range of pure and mixed input states we demonstrate via simulations that our method produces functionally equivalent reconstructed states to that of traditional methods with the added benefit that expensive computations are front-loaded with our system. Further, by training our system with measurement results that include simulated noise sources we are able to demonstrate a significantly enhanced average fidelity when compared to typical reconstruction methods. These enhancements in average fidelity are also shown to persist when we consider state reconstruction from partial tomography data where several measurements are missing. We anticipate that the present results combining the fields of machine intelligence and quantum state estimation will greatly improve and speed up tomography-based quantum experiments.
ZRLTAWVJ
webpage
Saved 2026-05-27
A few weeks ago my 5-year-old and I tried playing Cat Crimes, a puzzle game in which you work out which of your cats ate your shoes. We had a wonderful time - for about 20 minutes.
UG7WHFJT
webpage
Saved 2026-05-27
“What language is the Coffeescript compiler written in?” I asked my brilliant, witty and insightful colleague Grzegorz Kossakowski. “The Coffeescript compiler is written in Coffeescript,” he replied. Since this statement is clearly stupid, I denounced him as a liar and an idiot and took back all of the nice things I had said about him.
DB2QLAFY
webpage
Jimmy Donaldson
Saved 2026-05-27
2EZFPUPV
webpage
Saved 2026-05-27
Q7N83TNK
webpage
Saved 2026-05-26
9LRPBQP4
journalArticle
John A. Smolin, Jay M. Gambetta, Graeme Smith
2012 · Physical Review Letters
Saved 2026-05-26
We provide an efficient method for computing the maximum likelihood mixed quantum state (with density matrix $ρ$) given a set of measurement outcome in a complete orthonormal operator basis subject to Gaussian noise. Our method works by first changing basis yielding a candidate density matrix $μ$ which may have nonphysical (negative) eigenvalues, and then finding the nearest physical state under the 2-norm. Our algorithm takes at worst $O(d^4)$ for the basis change plus $O(d^3)$ for finding $ρ$ where $d$ is the dimension of the quantum state. In the special case where the measurement basis is strings of Pauli operators, the basis change takes only $O(d^3)$ as well. The workhorse of the algorithm is a new linear-time method for finding the closest probability distribution (in Euclidean distance) to a set of real numbers summing to one.
FZDJFKH6
blogPost
2026
Saved 2026-05-26
A lot of people seem convinced that the point of AI coding is to write low-quality code as fast as possible. Spew out barely-passable slop, open massive PRs, and merge them unvetted. Ship it! But t…
Z3P3AP27
webpage
Substack
2026
Saved 2026-05-25
When I was younger, I severally underestimated how valuable it is to compound, to focus in on a few things and do them well. There are things you can only discover 10,000 hours into a conversation with your best friend, experiences you can only have two decades into honing a craft or living in the same place. It takes dedication to not become blasé to the things we have around us, to maintain the serious attention and dedication necessary to discover new levels to something. But those levels are, in many ways, more rewarding than the rewards of doing new stuff (as fun as that is, and as important as it is to have that as part of the mix).
FE69FX39
book
On Killing
David Grossman
Saved 2026-05-25
47ZBS68H
webpage
Amitav Krishna
2026
Saved 2026-05-25
Hello! Quick throwaway post here. I was reading Normative Uncertainty by William MacAskill(Effective Altruism man) when I was flashbanged by the math. After some chatting with claude, apparently this is what makes analytic philosophy notable, as opposed to continental philosophy which is focused on society, culture, etc. What they don’t learn about apparently is analysis, which unlike logic, deals with continuous things.
WZM6DCFT
webpage
Michael Nielsen
Saved 2026-05-25
36XHI7LC
webpage
Michael Nielsen, Patrick Collison
2024
Saved 2026-05-25
AMEEZVJJ
book
The Scaling Book
Saved 2026-05-24
3CTDH87F
webpage
2024
Saved 2026-05-23
Mac's Tech Blog
JB3JIR8E
book
Hackauthor²
Saved 2026-05-23
Painlessly quit pornography immediately, without willpower or any sense of deprivation or sacrifice.
4M5Z7SIE
webpage
Saved 2026-05-22
VGHRGSXA
webpage
Saved 2026-05-22
Math Every Day Stevey's Drunken Blog Rants™
PURATK4W
webpage
Michael Nielsen
Saved 2026-05-22
ECX4IAXG
webpage
Substack
Saved 2026-05-22
and now I'm a different person
8958FSTV
journalArticle
Ethan Perez
2024
Saved 2026-05-22
TLDR: I’ve collected some tips for research that I’ve given to other people and/or used myself, which have sped things up and helped put people in th…
V2EYIIWT
webpage
2025
Saved 2026-05-21
Different forms of attention mechanisms, used in modern diffusion models.
YC7RGXQT
webpage
Substack
Saved 2026-05-21
You Don't Have Writers' Block; Just Keep Going
7KAMEAJW
journalArticle
Benjamin Schumacher, M. A. Nielsen
1996 · Physical Review A
Saved 2026-05-21
MN4X8L9R
journalArticle
A Spline Theory of Deep Networks
Randall Balestriero, Richard G Baraniuk
Saved 2026-05-21
We build a rigorous bridge between deep networks (DNs) and approximation theory via spline functions and operators. Our key result is that a large class of DNs can be written as a composition of max-affine spline operators (MASOs), which provide a powerful portal through which to view and analyze their inner workings. For instance, conditioned on the input signal, the output of a MASO DN can be written as a simple affine transformation of the input. This implies that a DN constructs a set of signal-dependent, class-specific templates against which the signal is compared via a simple inner product; we explore the links to the classical theory of optimal classification via matched filters and the effects of data memorization. Going further, we propose a simple penalty term that can be added to the cost function of any DN learning algorithm to force the templates to be orthogonal with each other; this leads to significantly improved classification performance and reduced overfitting with no change to the DN architecture. The spline partition of the input signal space opens up a new geometric avenue to study how DNs organize signals in a hierarchical fashion. As an application, we develop and validate a new distance metric for signals that quantifies the difference between their partition encodings.
BEHT77MJ
preprint
Thomas Fischbacher, Iulia M. Comsa, Krzysztof Potempa, Moritz Firsching, Luca Versari, Jyrki Alakuijala
2020
Saved 2026-05-21
We present a novel machine learning architecture that uses the exponential of a single input-dependent matrix as its only nonlinearity. The mathematical simplicity of this architecture allows a detailed analysis of its behaviour, providing robustness guarantees via Lipschitz bounds. Despite its simplicity, a single matrix exponential layer already provides universal approximation properties and can learn fundamental functions of the input, such as periodic functions or multivariate polynomials. This architecture outperforms other general-purpose architectures on benchmark problems, including CIFAR-10, using substantially fewer parameters.
X42MRP5L
preprint
Luke Darlow, Ciaran Regan, Sebastian Risi, Jeffrey Seely, Llion Jones
2025
Saved 2026-05-21
Biological brains demonstrate complex neural activity, where neural dynamics are critical to how brains process information. Most artificial neural networks ignore the complexity of individual neurons . We challenge that paradigm. By incorporating neuron-level processing and synchronization, we reintroduce neural timing as a foundational element. We present the Continuous Thought Machine (CTM), a model designed to leverage neural dynamics as its core representation. The CTM has two innovations: (1) neuron-level temporal processing, where each neuron uses unique weight parameters to process incoming histories; and (2) neural synchronization as a latent representation. The CTM aims to strike a balance between neuron abstractions and biological realism. It operates at a level of abstraction that effectively captures essential temporal dynamics while remaining computationally tractable. We demonstrate the CTM’s performance and versatility across a range of tasks, including solving 2D mazes, ImageNet1K classification, parity computation, and more. Beyond displaying rich internal representations and offering a natural avenue for interpretation owing to its internal process, the CTM is able to perform tasks that require complex sequential reasoning. The CTM can also leverage adaptive compute, where it can stop earlier for simpler tasks, or keep computing when faced with more challenging instances. The goal of this work is to share the CTM and its associated innovations, rather than pushing for new state-of-the-art results. To that end, we believe the CTM represents a significant step toward developing more biologically plausible and powerful artificial intelligence systems. We provide an accompanying interactive online demonstration and an extended technical report.
HBX85V6G
webpage
Sakana AI
2026
Saved 2026-05-21
新しいBusiness Intelligenceへ:Ultra Deep Researchアシスタント「Sakana Marlin」βテスト開始
RAZRJ4GH
blogPost
Mattan Griffel
2018
Saved 2026-05-20
5 rules for good email etiquette
JZRIJKC8
webpage
Saved 2026-05-20
RGDWELX2
webpage
Saved 2026-05-20
Parenthetically Speaking: Articles by Shriram Krishnamurthi
IAL827EV
webpage
Cohere Labs Community
2026
Saved 2026-05-20
A community lead reflects on three years of learning, research, and building programs inside the Cohere Labs Open Science Community.
ZZ5MWM73
webpage
William MacAskill
Saved 2026-05-20
7UVC593T
webpage
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis et al.
2020
Saved 2026-05-20
Large pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks. However, their ability to access and precisely manipulate knowledge is still limited, and hence on knowledge-intensive tasks, their performance lags behind task-specific architectures. Additionally, providing provenance for their decisions and updating their world knowledge remain open research problems. Pre-trained models with a differentiable access mechanism to explicit non-parametric memory can overcome this issue, but have so far been only investigated for extractive downstream tasks. We explore a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation. We introduce RAG models where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. We compare two RAG formulations, one which conditions on the same retrieved passages across the whole generated sequence, the other can use different passages per token. We fine-tune and evaluate our models on a wide range of knowledge-intensive NLP tasks and set the state-of-the-art on three open domain QA tasks, outperforming parametric seq2seq models and task-specific retrieve-and-extract architectures. For language generation tasks, we find that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline.
GFVWKXG5
webpage
Saved 2026-05-20
IHLDKNFY
webpage
Saved 2026-05-20
7W53YSF6
webpage
Saved 2026-05-20
What is the reason for proliferation of DSLs in the last year?
86K27KWP
webpage
Saved 2026-05-20
Form and track positive lasting habits built with 💙 by me - powered by org 🦄 Why Keeping habits accessible and trackable has helped me form good hab...
5C3D28RH
journalArticle
Journal of Natural Science and Exploration
Ekagrata Bahadur, Amrit Nath Thulal
Saved 2026-05-20
A Natural language processing (NLP) has increased the interest in genetic algorithm (GA) due to their skills in solving complex optimization problems with extensive research on the use of genetic algorithms in NLP projects has been presented in this paper. First, we present the basic concepts behind genetic algorithms and their relevance to natural language processing. Then, we explore various applications of natural language processing (NLP) that use genetic algorithms, including text classification, sentiment analysis, machine translation, summarization, and question-answering systems. We examine the advantages and disadvantages of genetic algorithm applications in natural language processing by comparing their performance with traditional and modern approaches and discuss the factors influencing their effectiveness. Furthermore, we explore recent advancements, modifications, and hybridizations of Genetic Algorithms tailored to NLP tasks. Finally, we discuss the challenges and future directions in leveraging Genetic Algorithms for enhancing NLP technologies.
B7XFUCZM
webpage
Saved 2026-05-20
Confluent is building the foundational platform for data in motion so any organization can innovate and win in a digital-first world.
CT4L46N4
webpage
2026
Saved 2026-05-20
A few days ago, I posted about a personal project that I've been working on for the last few weeks: Mr. Chatterbox, a chatbot trained from scratch on Victorian-era literature. I have to admit, I was totally blown away (and a little frightened!) by the reception. Whenever I post about
G3ZFHKRY
webpage
Saved 2026-05-20
Media over QUIC: There are ways to do voice AI without being traumatized by WebRTC.
MJNCCFAG
webpage
Saved 2026-05-19
UWSSKBJ6
webpage
Saved 2026-05-19
COS568 Systems and Machine Learning (Spring 2025) Programming Assignments Network pruning Distributed training of a language model Project Learned index Tentative Syllabus Dates Presenters Topics & Main Papers Related Papers Events 1/31 Kai Li Dr. Jeff Dean & Dr. Amin Vahdat (Goo...
R52LUFMW
webpage
Vlad Feinberg
2025
Saved 2026-05-19
Vlad's Blog
QT2FVMYB
webpage
Vlad Feinberg
2026
Saved 2026-05-18
Vlad's Blog
3HT5CZX7
webpage
Saved 2026-05-18
How the rising star of podcasting eschews the breadth v. depth dilemma and is quickly becoming known as ‘the new Lex Fridman.’
ADWIMTN9
newspaperArticle
Katya Ungerman
2026 · The New York Times
Saved 2026-05-18
One of the oldest and most durable features of human experience is re-emerging.
DW9GGNSS
webpage
Saved 2026-05-17
We launched the Datasette Cloud blog today. The Datasette Cloud site itself is a Django app - it uses Django and PostgreSQL to manage accounts, teams and soon billing and payments, then launches dedicated containers running Datasette for each customer.
UTCVFFCG
webpage
Saved 2026-05-17
Y5IKS6HK
webpage
Saved 2026-05-17
Nextpad++ feels like a fever dream. Like what Mac apps would be if the Nazis had won WWII.
5DRBECLG
webpage
Saved 2026-05-17
It’s not even a feature. It’s just technology.
QUWPVPZS
webpage
Simon Willison
Saved 2026-05-17
François Chollet is the co-founder of the ARC Prize and had advanced access to today's o3 results. His article here is the most insightful coverage I've seen of o3, going …
W3KQRHHN
webpage
Simon Willison
Saved 2026-05-17
You should start a blog. Having your own little corner of the internet is good for the soul! But what should you write about? It’s easy to get hung up …
NBI5SJ4C
webpage
Saved 2026-05-17
POSSE is an abbreviation for Publish (on your) Own Site, Syndicate Elsewhere, the practice of posting content on your own site first, then publishing copies or sharing links to third parties (like social media silos) with original post links to provide viewers a path to directly interacting with your content.
RIKWAVLC
webpage
Simon Willison
Saved 2026-05-17
I started running a basic link blog on this domain back in November 2003—publishing links (which I called “blogmarks”) with a title, URL, short snippet of commentary and a “via” …
469NQS3D
webpage
Jason Koebler ·
2026
Saved 2026-05-17
AI writing is impossible to avoid, is making everything sound the same, and is driving us crazy.
PEJMAT5S
blogPost
Zohar Atkins
2026
Saved 2026-05-17
Walter Benjamin and Rabbi Levi Yitzchak of Berditchev on Translation
VFV2ZZXL
webpage
Leo Tolstoy
Saved 2026-05-17
ZKJLDT7I
webpage
2025
Saved 2026-05-17
Notes on applied AI engineering, machine learning, and data science.
3H73EWEN
blogPost
Malmesbury
2026
Saved 2026-05-17
Recommended soundtrack for this post
KWTDAYJB
preprint
Haoyi Zhu, Haozhe Liu, Yuyang Zhao, Tian Ye, Junsong Chen, Jincheng Yu, Tong He, Song Han et al.
2026
Saved 2026-05-17
We introduce SANA-WM, an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p, minute-scale videos with precise camera control. SANA-WM achieves visual quality comparable to large-scale industrial baselines such as LingBot-World and HY-WorldPlay, while significantly improving efficiency. Four core designs drive our architecture: (1) Hybrid Linear Attention combines frame-wise Gated DeltaNet (GDN) with softmax attention for memory-efficient long-context modeling. (2) Dual-Branch Camera Control ensures precise 6-DoF trajectory adherence. (3) Two-Stage Generation Pipeline applies a long-video refiner to stage-1 outputs, improving quality and consistency across sequences. (4) Robust Annotation Pipeline extracts accurate metric-scale 6-DoF camera poses from public videos to yield high-quality, spatiotemporally consistent action labels. Driven by these designs, SANA-WMdemonstrates remarkable efficiency across data, training compute, and inference hardware: it uses only $\sim$213K public video clips with metric-scale pose supervision, completes training in 15 days on 64 H100s, and generates each 60s clip on a single GPU; its distilled variant can be deployed on a single RTX 5090 with NVFP4 quantization to denoise a 60s 720p clip in 34s. On our one-minute world-model benchmark, SANA-WM demonstrates stronger action-following accuracy than prior open-source baselines and achieves comparable visual quality at $36\times$ higher throughput for scalable world modeling.
WB7KQR3H
webpage
Saved 2026-05-16
PYV9I8V9
preprint
Hao Tang, Darren Key, Kevin Ellis
2024
Saved 2026-05-16
We give a model-based agent that builds a Python program representing its knowledge of the world based on its interactions with the environment. The world model tries to explain its interactions, while also being optimistic about what reward it can achieve. We define this optimism as a logical constraint between a program and a planner. We study our agent on gridworlds, and on task planning, finding our approach is more sample-efficient compared to deep RL, more compute-efficient compared to ReAct-style agents, and that it can transfer its knowledge across environments by editing its code.
2EUFFSSC
webpage
Saved 2026-05-15
Personal Site
NSG8IBQI
forumPost
Lucky_Wrap
2019
Saved 2026-05-15
J5KZDL4H
webpage
Saved 2026-05-15
3 track album
6WQ4QICZ
journalArticle
Junhyuk Oh, Gregory Farquhar, Iurii Kemaev, Dan A. Calian, Matteo Hessel, Luisa Zintgraf, Satinder Singh, Hado van Hasselt et al.
2025 · Nature · Nature Publishing Group
Saved 2026-05-15
Humans and other animals use powerful reinforcement learning (RL) mechanisms that have been discovered by evolution over many generations of trial and error. By contrast, artificial agents typically learn using handcrafted learning rules. Despite decades of interest, the goal of autonomously discovering powerful RL algorithms has proven to be elusive1–6. Here we show that it is possible for machines to discover a state-of-the-art RL rule that outperforms manually designed rules. This was achieved by meta-learning from the cumulative experiences of a population of agents across a large number of complex environments. Specifically, our method discovers the RL rule by which the agent’s policy and predictions are updated. In our large-scale experiments, the discovered rule surpassed all existing rules on the well-established Atari benchmark and outperformed a number of state-of-the-art RL algorithms on challenging benchmarks that it had not seen during discovery. Our findings suggest that the RL algorithms required for advanced artificial intelligence may soon be automatically discovered from the experiences of agents, rather than manually designed.
JESHI4F9
webpage
Thinking Machines Lab
Saved 2026-05-15
Interaction models move beyond turn-based AI interfaces by handling multimodal, real-time collaboration natively across audio, video, and text.
4PLTSZBD
book
Stoner
John Williams
2003 · New York Review Books
Saved 2026-05-15
UVLHGE6F
forumPost
Adventurous-Math-322
2026
Saved 2026-05-15
784VK76U
webpage
Dr Werner Vogels- https://www.allthingsdistributed.com
2026
Saved 2026-05-14
The subtle inventiveness that reduced cold start setup from seconds to 200μs.
5QT54B2H
journalArticle
Kent C. Berridge, Terry E. Robinson
2016 · American Psychologist
Saved 2026-05-13
Rewards are both ‘liked’ and ‘wanted’, and those two words seem almost interchangeable. However, the brain circuitry that mediates the psychological process of ‘wanting’ a particular reward is dissociable from circuitry that mediates the degree to which it is ‘liked’. Incentive salience or ‘wanting’, a form of motivation, is generated by large and robust neural systems that include mesolimbic dopamine. By comparison, ‘liking’, or the actual pleasurable impact of reward consumption, is mediated by smaller and fragile neural systems, and is not dependent on dopamine. The incentive-sensitization theory posits the essence of drug addiction to be excessive amplification specifically of psychological ‘wanting’, especially triggered by cues, without necessarily an amplification of ‘liking’. This is due to long-lasting changes in dopamine-related motivation systems of susceptible individuals, called neural sensitization. A quarter-century after its proposal, evidence has continued to grow in support the incentive-sensitization theory. Further, its scope is now expanding to include diverse behavioral addictions and other psychopathologies.
6P6NJEQD
webpage
Viola Zhou
2026
Saved 2026-05-13
Chinese-born tech workers have fueled Silicon Valley for decades. In the AI era, they're superstars.
5D57K3UE
bookSection
Aslak Tveito, Are Magnus Bruaset, Olav Lysne, J. F. Kaiser
2010 · Simula Research Laboratory · Springer Berlin Heidelberg
Saved 2026-05-13
3DQWLHSS
webpage
Saved 2026-05-12
JHQ2RHTI
preprint
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
Fan Long
Saved 2026-05-12
KKFWFEG5
preprint
Tara Saba, Anne Ouyang, Xujie Si, Fan Long
2026
Saved 2026-05-12
High-performance GPU kernels are critical to modern machine learning systems, yet developing efficient implementations remains a challenging, expert-driven process due to the tight coupling between algorithmic structure, memory hierarchy usage, and hardware-specific optimizations. Recent work has explored using large language models (LLMs) to generate GPU kernels automatically, but generated implementations often struggle to maintain correctness and achieve competitive performance across iterative refinements. We present CuTeGen, an agentic framework for automated generation and optimization of GPU kernels that treats kernel development as a structured generate--test--refine workflow. Unlike approaches that rely on one-shot generation or large-scale search over candidate implementations, CuTeGen focuses on progressive refinement of a single evolving kernel through execution-based validation, structured debugging, and staged optimization. A key design choice is to generate kernels using the CuTe abstraction layer, which exposes performance-critical structures such as tiling and data movement while providing a more stable representation for iterative modification. To guide performance improvement, CuTeGen incorporates workload-aware optimization prompts and delayed integration of profiling feedback. Experimental results on matrix multiplication and activation workloads demonstrate that the framework produces functionally correct kernels and achieves competitive performance relative to optimized library implementations.
NZM98V5D
conferencePaper
Fan Long, Martin Rinard
2016 · Proceedings of the 43rd Annual ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages · ACM
Saved 2026-05-12
HSYCIZGZ
preprint
Yuhui Li, Fangyun Wei, Chao Zhang, Hongyang Zhang
2025
Saved 2026-05-12
The sequential nature of modern LLMs makes them expensive and slow, and speculative sampling has proven to be an effective solution to this problem. Methods like EAGLE perform autoregression at the feature level, reusing top-layer features from the target model to achieve better results than vanilla speculative sampling. A growing trend in the LLM community is scaling up training data to improve model intelligence without increasing inference costs. However, we observe that scaling up data provides limited improvements for EAGLE. We identify that this limitation arises from EAGLE's feature prediction constraints. In this paper, we introduce EAGLE-3, which abandons feature prediction in favor of direct token prediction and replaces reliance on top-layer features with multi-layer feature fusion via a technique named training-time test. These improvements significantly enhance performance and enable the draft model to fully benefit from scaling up training data. Our experiments include both chat models and reasoning models, evaluated on five tasks. The results show that EAGLE-3 achieves a speedup ratio up to 6.5x, with about 1.4x improvement over EAGLE-2. In the SGLang framework, EAGLE-3 achieves a 1.38x throughput improvement at a batch size of 64. The code is available at https://github.com/SafeAILab/EAGLE.
JVC7QYW3
preprint
Hongyang Zhang
Saved 2026-05-12
NYXB3PTG
blogPost
Matt Giaro
2025
Saved 2026-05-12
This app is the backbone of my 6-figure writing business
LFJAIATC
webpage
Saved 2026-05-12
2CEM2ZK5
journalArticle
From Worm to Human: Scaling Brain Emulation
Isaak Freeman
Saved 2026-05-11
Machine learning models are rapidly approaching or surpassing human performance on many metrics. In comparison, neuroscience is progressing at a slow pace. To reach the research velocity possible in software paradigms, we highlight a potential path to high-quality emulations of the brains of key model organisms such as the roundworm Caenorhabditis elegans, the zebrafish Danio rerio and the mouse Mus musculus.
AGRC2GVW
blogPost
John Friedman
2026
Saved 2026-05-11
Scaling data processing, starting with the SEC corpus.
3CAUY7QL
blogPost
John Friedman
2026
Saved 2026-05-11
Addressing a recent, popular misconception
LF43M5BV
blogPost
John Friedman
2025
Saved 2026-05-11
I had fun.
NI59DF55
webpage
2026
Saved 2026-05-11
from a European fresh off the boat
XGLETZKS
webpage
2004
Saved 2026-05-11
Ted Leung noted the discussion that Werner and I have been having, and observed that we should consider Rob Pike’s (in)famous polemic, “Systems Software Research is Irrelevant.” I should say that I broadly agree with most of Pike’s conclusions – and academic systems software research has seemed increasingly irrelevant in the last five years. That said, I think that what Pike characterizes as “systems research” is far too skewed to the interface to the system – which (tautologically) is but the periphery of the larger system. In my opinion, “systems research” should focus not on the interface of the system, but rather its guts: those hidden Rube Goldberg-esque innards that are rife with old assumptions and unintended consequences. Pike would perhaps dismiss the study of these innards as “phenomenology”, but I would counter that understanding phenomena is a prerequisite to understanding larger systemic truths. Of course, the problem to date has been that much systems research has not been able to completely understand phenomena – the research has often consisted merely of characterizing it.
MNH47WSR
webpage
2025
Saved 2026-05-11
Note: Thie was co-authored with Steve Tuck, and originally appeared on the Oxide blog. We don’t want to bury the lede: we have raised a $100M Series B, led by a new strategic partner in USIT with participation from all existing Oxide investors. To put that number in perspective: over the nearly six year lifetime of the company, we have raised $89M; our $100M Series B more than doubles our total capital raised to date — and positions us to make Oxide the generational company that we have always aspired it to be.
H2RHMNUZ
webpage
2016
Saved 2026-05-11
Last Tuesday, several months of preparation came to fruition in the inaugural Systems We Love. You never know what’s going to happen the first time you get a new kind of conference together (especially one as broad as this one!) but it was, in a word, amazing. The content was absolutely outstanding, with attendee after attendee praising the uniformly high quality. (For guided tours, check out both Ozan Onay’s excellent exegesis and David Cassel’s thorough New Stack story – and don’t miss Sarah Huffman’s incredible illustrations!) It was such a great conference that many were asking about when we would do it again – and there is already interest in replicating it elsewhere. As an engineer, this makes me slightly nervous as I believe that success often teaches you nothing: luck becomes difficult to differentiate from design. But at the risk of taunting the conference gods with the arrogance of a puny mortal, here’s some stuff I do think we did right:
RQ588BYF
webpage
2025
Saved 2026-05-11
USENIX made the decision this week to discontinue its flagship Annual Technical Conference. When USENIX was started in 1975 — before the Internet, really — conferences were the fastest vector for practitioners to formally share their ideas, and USENIX ATC flourished. Speaking for myself, I came up lionizing ATC: I was an undergraduate in the early 1990s, and programs like the USENIX Summer 1994 conference felt like Renaissance-era Florence for systems practitioners.
XUI6NAKY
journalArticle
Systems Software Research is Irrelevant
Rob Pike, Bell Labs
Saved 2026-05-11
V57EMEKD
webpage
2019
Saved 2026-05-10
Derek Sivers official site. Thoughts on philosophy, culture, self-improvement. Author of Useful Not True, How to Live, Hell Yeah or No, Anything You Want.
H45DILSI
webpage
2018
Saved 2026-05-10
Derek Sivers official site. Thoughts on philosophy, culture, self-improvement. Author of Useful Not True, How to Live, Hell Yeah or No, Anything You Want.
CXPT4MHB
webpage
2026
Saved 2026-05-10
Derek Sivers official site. Thoughts on philosophy, culture, self-improvement. Author of Useful Not True, How to Live, Hell Yeah or No, Anything You Want.
JMCH5W9J
webpage
Saved 2026-05-10
LCHH25TW
webpage
2023
Saved 2026-05-10
AXEE9BA3
webpage
Saved 2026-05-10
8HIJUU8Z
magazineArticle
James Somers
2018 · The New Yorker
Saved 2026-05-09
Coding together at the same computer, Jeff Dean and Sanjay Ghemawat changed the course of the company—and the Internet.
JPPEJZ3L
webpage
Christopher Olah
Saved 2026-05-09
Instead of asking “Is university good?”, ask “Do I have something more compelling to do?”. Instead of “Should I do a PhD?”, ask “Where can I find the best environment to grow as a researcher?”.
TC3I77TI
webpage
2026
Saved 2026-05-09
Many founders choose the factory, and they choose it for understandable reasons: safety in numbers, a clear set of instructions, the comfort of a recognizable path. But a factory has one purpose: to produce more of the same. If you are making Fords, more of the same is good. But
4SJP2TIP
webpage
Saved 2026-05-08
RDHGQE3C
journalArticle
Neel Nanda
2025
Saved 2026-05-08
Last updated Sept 2 2025 • Note - if you want to pursue a career in this kind of research, apply to my MATS stream! Due Dec 23 …
MD92Z73A
conferencePaper
Andrew Boutros, Eriko Nurvitadhi, Rui Ma, Sergey Gribok, Zhipeng Zhao, James C. Hoe, Vaughn Betz, Martin Langhammer
2020 · 2020 International Conference on Field-Programmable Technology (ICFPT) · IEEE
Saved 2026-05-07
The growing importance and compute demands of artificial intelligence (AI) have led to the emergence of domainoptimized hardware platforms. For example, Nvidia GPUs introduced specialized tensor cores for matrix operations to speed up deep learning (DL) computation, resulting in very high peak throughput up to 130 int8 TOPS in the T4 GPU. Recently, Intel introduced its first AI-optimized 14nm FPGA, the Stratix 10 NX, with in-fabric AI tensor blocks that offer estimated peak performance up to 143 int8 TOPS, comparable to 12nm GPUs. However, what matters in practice is not the peak performance but the actual achievable performance on target workloads. This depends mainly on the utilization of the tensor units, and the system-level overheads to send data to/from the accelerator.
E37N55L2
journalArticle
Pouya Kananian, Arnesh Sujanani, Seyed Majid Zahedi
2026 · Journal of Artificial Intelligence Research
Saved 2026-05-07
We study the fair and truthful allocation of m divisible public items among n agents, each with distinct preferences for the items. To aggregate agents’ preferences fairly, we focus on finding a core solution. For divisible items, a core solution always exists and can be calculated by maximizing the Nash welfare objective. However, such a solution is easily manipulated; agents might have incentives to misreport their preferences. To mitigate this, the current state-of-the-art finds an approximate core solution with high probability while ensuring approximate truthfulness. However, this approach has two main limitations. First, due to several approximations, the approximation error in the core could grow with n, resulting in a non-asymptotic core solution. This limitation is particularly significant as public-good allocation mechanisms are frequently applied in scenarios involving a large number of agents, such as the allocation of public tax funds for municipal projects. Second, implementing the current approach for practical applications proves to be a highly nontrivial task. To address these limitations, we introduce PPGA, a (differentially) Private Public-Good Allocation algorithm, and show that it attains asymptotic truthfulness and finds an asymptotic core solution with high probability. Additionally, to demonstrate the practical applicability of our algorithm, we implement PPGA and empirically study its properties using municipal participatory budgeting data.
V3H2JTRM
preprint
Jiashu Zhang, Zihan Pan, Molly, Xu, Khuzaima Daudjee, Sihang Liu
2025
Saved 2026-05-07
The occurrence of bubbles in pipeline parallelism is an inherent limitation that can account for more than 40% of the large language model (LLM) training time and is one of the main reasons for the underutilization of GPU resources in LLM training. Harvesting these bubbles for GPU side tasks can increase resource utilization and reduce training costs but comes with challenges. First, because bubbles are discontinuous with various shapes, programming side tasks becomes difficult while requiring excessive engineering effort. Second, a side task can compete with pipeline training for GPU resources and incur significant overhead. To address these challenges, we propose FreeRide, a system designed to harvest bubbles in pipeline parallelism for side tasks. FreeRide provides programmers with interfaces to implement side tasks easily, manages bubbles and side tasks during pipeline training, and controls access to GPU resources by side tasks to reduce overhead. We demonstrate that FreeRide achieves 7.8% average cost savings with a negligible overhead of about 1% in training LLMs while serving model training, graph analytics, and image processing side tasks.
9XLB5W5V
preprint
Desen Sun, Shuncheng Jie, Sihang Liu
2026
Saved 2026-05-07
Diffusion models are a powerful class of generative models that produce images and other content from user prompts, but they are computationally intensive. To mitigate this cost, recent academic and industry work has adopted approximate caching, which reuses intermediate states from similar prompts in a cache. While efficient, this optimization introduces new security risks by breaking isolation among users. This paper provides a comprehensive assessment of the security vulnerabilities introduced by approximate caching. First, we demonstrate a remote covert channel established with the approximate cache, where a sender injects prompts with special keywords into the cache system and a receiver can recover that even after days, to exchange information. Second, we introduce a prompt stealing attack using the approximate cache, where an attacker can recover existing cached prompts from hits. Finally, we introduce a poisoning attack that embeds the attacker's logos into the previously stolen prompt, leading to unexpected logo rendering for the requests that hit the poisoned cache prompts. These attacks are all performed remotely through the serving system, demonstrating severe security vulnerabilities in approximate caching. The code for this work is available.
7FSRP4ID
conferencePaper
Yaoyao Ding, Bohan Hou, Xiao Zhang, Allan Lin, Tianqi Chen, Cody Hao Yu, Yida Wang, Gennady Pekhimenko
2026 · Proceedings of the 31st ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 · ACM
Saved 2026-05-07
YV7VS5YG
webpage
Julia Evans
2019
Saved 2026-05-07
Get your work recognized: write a brag document
GYJ86NLY
preprint
Mohammad Asadi, Jack W. O'Sullivan, Fang Cao, Tahoura Nedaee, Kamyar Rajabalifardi, Fei-Fei Li, Ehsan Adeli, Euan Ashley
2026
Saved 2026-05-06
Multimodal AI systems have achieved remarkable performance across a broad range of real-world tasks, yet the mechanisms underlying visual–language reasoning remain surprisingly poorly understood. We report three findings that challenge prevailing assumptions about how these systems process and integrate visual information. First, Frontier models readily generate detailed image descriptions and elaborate reasoning traces, including pathology-biased clinical findings, for images never provided; we term this phenomenon mirage reasoning. Second, without any image input, models also attain strikingly high scores across general and medical multimodal benchmarks, bringing into question their utility and design. In the most extreme case, our model achieved the top rank on a standard chest Xray question-answering benchmark without access to any images. Third, when models were explicitly instructed to guess answers without image access, rather than being implicitly prompted to assume images were present, performance declined markedly. Explicit guessing appears to engage a more conservative response regime, in contrast to the mirage regime in which models behave as though images have been provided. These findings expose fundamental vulnerabilities in how visual–language models reason and are evaluated, pointing to an urgent need for private benchmarks that eliminate textual cues enabling non-visual inference, particularly in medical contexts where miscalibrated AI carries the greatest consequence. We introduce B-Clean as a principled solution for fair, vision-grounded evaluation of multimodal AI systems.
RUFLLDHM
webpage
Saved 2026-05-06
6T26ANB9
webpage
Saved 2026-05-06
CHJX9UQT
webpage
Saved 2026-05-06
3NYJN9Z8
webpage
Julia Evans
2016
Saved 2026-05-06
How to ask good questions
9S7RH8QG
journalArticle
Curry-Howard Correspondence
Giselle Reis
Saved 2026-05-06
36TWY7KW
webpage
Saved 2026-05-06
FFZE8LCK
journalArticle
Manifesto of the Communist Party
Karl Marx
Saved 2026-05-04
CIZWPSMD
preprint
George Morgulis, John Hewitt
2026
Saved 2026-05-04
Subliminal learning describes a student language model inheriting a behavioral bias by fine-tuning on seemingly innocuous data generated by a biased teacher model. Prior work has begun to characterize this phenomenon but leaves open questions about the scope of signals it can transfer, the mechanisms that explain it, and the precision with which a bias can be encoded by seemingly unrelated data. We tackle all three problems by introducing subliminal steering, a variant of subliminal learning in which the teacher’s bias is implemented not via a system prompt, as in prior work, but through a steering vector trained to maximize the likelihood of a set of target samples. First, we show that subliminal steering transfers complex multi-word biases, whereas prior work focused on single-word preferences—demonstrating a large scope of subliminally transferrable signals. Second, we provide mechanistic evidence that subliminal learning transfers not only the target behavioral bias, but also the steering vector itself, localized to the layers at which the teacher was steered. Finally, we show that the bias is encoded with surprising precision. We train a new steering vector directly on the subliminally-laden dataset and find that it attains high cosine similarity with the original vector.
RGXBTLDN
webpage
Elon Litman
2025
Saved 2026-05-04
How the search for a simple approximation revealed a beautiful, exact solution.
DH5ZUCEJ
blogPost
Laura Schroeder
2026
Saved 2026-05-04
on clarity, endurance, and putting yourself out there
LLP359S6
blogPost
Tea
2010
Saved 2026-05-03
AF7AF4WW
preprint
Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J. Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q. Tran et al.
2024
Saved 2026-05-02
More than three billion years of evolution have produced an image of biology encoded into the space of natural proteins. Here we show that language models trained on tokens generated by evolution can act as evolutionary simulators to generate functional proteins that are far away from known proteins. We present ESM3, a frontier multimodal generative language model that reasons over the sequence, structure, and function of proteins. ESM3 can follow complex prompts combining its modalities and is highly responsive to biological alignment. We have prompted ESM3 to generate fluorescent proteins with a chain of thought. Among the generations that we synthesized, we found a bright fluorescent protein at far distance (58% identity) from known fluorescent proteins. Similarly distant natural fluorescent proteins are separated by over five hundred million years of evolution.
LRKKSZSU
forumPost
Yacine Mahdid [@yacinelearning]
2025
Saved 2026-05-01
P7XNVSX2
blogPost
Derrick Harris, Matt Bornstein, Guido Appenzeller
2023
Saved 2026-05-01
A curated list of resources we’ve relied on to get smarter about modern AI, including generative AI, LLMs, and transformer models.
DAM7J2SN
webpage
Saved 2026-04-29
XKRN9D3D
webpage
Ross Benes
Saved 2026-04-28
XI3AXMD9
magazineArticle
· The Economist
Saved 2026-04-28
L5ENGT2K
webpage
Tim Richardson
Saved 2026-04-28
'Severe punishment' for online miscreants
6XAW6L2T
blogPost
Tyler Cowen
2026
Saved 2026-04-28
From MR commentator Sure: Generally such figures do not reside within the physicians’ office. On our side of the table we do some procedure with multiple specifications and generate some CPT code(s) (e.g. a lap cholycystectomy is 47562, add on a common bile duct exploration and it becomes a 47564, and if you just do […]
UUDZIXDS
webpage
Kelsey Piper
2026
Saved 2026-04-28
AI only needs 150 words to identify me. What does that mean for you?
ZCECW6F6
webpage
2026
Saved 2026-04-28
AI can echolocate authors through their prose. Your digital fingerprint is at risk.
7U75TIEQ
webpage
Saved 2026-04-28
We create innovative outdoor products and experiences that fund sustainable poverty relief, move people to do good, and inspire adventure.
LP4T5FPQ
webpage
Saved 2026-04-28
Andela provides the human compute layer behind modern AI systems — training models,
 deploying AI-native engineers, and upskilling the teams that build them.
GT6Q8XAR
blogPost
Aaryan Harshith
2026
Saved 2026-04-27
And why reading it is the most meaningful challenge in biology
GZLB93BP
journalArticle
Semantic Self-Consistency: Enhancing Language Model Reasoning via Semantic Weighting
Tim Knappe, Ryan Li, Ayush Chauhan, Kaylee Chhua, Kevin Zhu, Sean O’Brien
Saved 2026-04-26
While large language models (LLMs) have rapidly improved their performance on a broad number of tasks, they still often fall short on reasoning tasks. As LLMs become more integrated in diverse real-world tasks, advancing their reasoning capabilities is crucial to their effectiveness in nuanced, complex problems. Wang et al. [33]’s self-consistency framework reveals that sampling multiple rationales before taking a majority vote reliably improves model performance across various closed-answer reasoning tasks. Standard methods based on this framework aggregate the final decisions of these rationales but fail to utilize the semantic information detailed in the step-by-step reasoning paths. Our work introduces semantic self-consistency, enhancing this approach by incorporating and analyzing both the reasoning paths of these rationales in addition to their final decisions before taking a majority vote. These methods not only improve the reliability of reasoning paths but also cause more robust performance on complex reasoning tasks.
VAMTKFWW
preprint
Timothy P. Lillicrap, Daniel Cownden, Douglas B. Tweed, Colin J. Akerman
2014
Saved 2026-04-26
The brain processes information through many layers of neurons. This deep architecture is representationally powerful, but it complicates learning by making it hard to identify the responsible neurons when a mistake is made. In machine learning, the backpropagation algorithm assigns blame to a neuron by computing exactly how it contributed to an error. To do this, it multiplies error signals by matrices consisting of all the synaptic weights on the neuron's axon and farther downstream. This operation requires a precisely choreographed transport of synaptic weight information, which is thought to be impossible in the brain. Here we present a surprisingly simple algorithm for deep learning, which assigns blame by multiplying error signals by random synaptic weights. We show that a network can learn to extract useful information from signals sent through these random feedback connections. In essence, the network learns to learn. We demonstrate that this new mechanism performs as quickly and accurately as backpropagation on a variety of problems and describe the principles which underlie its function. Our demonstration provides a plausible basis for how a neuron can be adapted using error signals generated at distal locations in the brain, and thus dispels long-held assumptions about the algorithmic constraints on learning in neural circuits.
RKVKZ7LK
blogPost
David Stutz
2024
Saved 2026-04-26
The decision to have a separate High School Project Track at NeurIPS 2024 has sparked quite some controversy, with many prominent AI researchers debating pros and cons and personal opinions, primarily on X/Twitter. Initially, I ignored this discussion, but eventually started thinking about it myself. Here are some of my thoughts.
IY87HZXJ
webpage
Zilan Qian
2025
Saved 2026-04-25
Life inside the "human sea attack"
53LZVIDP
journalArticle
Maison Clouâtré, Stefano Marano, Peter L. Falb, Moe Z. Win
2024 · IEEE Control Systems Letters
Saved 2026-04-24
This letter investigates parameter estimation in quantum systems that undergo dynamical evolution. Optimal control problems are formulated to maximize the information, about an unknown parameter, extracted by a given quantum measurement apparatus. This letter introduces the concept of “admissible controls”—control laws that do not depend on the unknown parameter they elicit. For scalar parameter estimation in unital quantum systems interrogated by binary measurements, this letter derives a necessary and sufficient condition on quantum measurement operators so that an information maximizing control law is admissible. When the admissibility condition is satisfied, it is shown that the resulting optimal control problem may be solved using well-established techniques.
BBLRCC9C
blogPost
Chamath Palihapitiya
2026
Saved 2026-04-24
2025 was, in some ways, a historic year for Social Capital. Numerically, we did well. Our portfolio took an important inflection upwards when NVIDIA licensed Groq for $20B. However...
3J65XLZ8
webpage
Erik Torenberg
2026
Saved 2026-04-24
America | Tech | Opinion | Culture | Charts
4CWU4XEN
webpage
Erik Torenberg
2026
Saved 2026-04-24
America | Tech | Opinion | Culture | Charts
TH4A97RU
journalArticle
John A. Smolin, Jay M. Gambetta, Graeme Smith
2012 · Physical Review Letters
Saved 2026-04-24
We provide an efficient method for computing the maximum likelihood mixed quantum state (with density matrix $ρ$) given a set of measurement outcome in a complete orthonormal operator basis subject to Gaussian noise. Our method works by first changing basis yielding a candidate density matrix $μ$ which may have nonphysical (negative) eigenvalues, and then finding the nearest physical state under the 2-norm. Our algorithm takes at worst $O(d^4)$ for the basis change plus $O(d^3)$ for finding $ρ$ where $d$ is the dimension of the quantum state. In the special case where the measurement basis is strings of Pauli operators, the basis change takes only $O(d^3)$ as well. The workhorse of the algorithm is a new linear-time method for finding the closest probability distribution (in Euclidean distance) to a set of real numbers summing to one.
329SGRQG
webpage
2026
Saved 2026-04-23
Introducing GPT-5.5, our smartest model yet—faster, more capable, and built for complex tasks like coding, research, and data analysis across tools.
UEXE69RH
webpage
2026
Saved 2026-04-23
We’ve redesigned our engineering interview process from the ground up.
Q2VCE2FT
blogPost
Chamath Palihapitiya
2026
Saved 2026-04-22
2025 was, in some ways, a historic year for Social Capital. Numerically, we did well. Our portfolio took an important inflection upwards when NVIDIA licensed Groq for $20B. However...
T3QPJUZT
blogPost
Tyler Cowen
2019
Saved 2026-04-22
Following up on my post a few days ago, about the value of deliberate practice for knowledge workers, a number of you asked me what form my practice takes.  A few of you were skeptical, but it is long since established that practice improves both your writing and your memory, so surely it can do […]
I4EH752N
webpage
Substack
Saved 2026-04-21
why the most respected engineers are pushing back on the coding agent hype cycle, and what the claude code argument actually reveals
KMRGEYTT
book
Quantum circuit fidelity estimation using machine learning
Avi Vadali, Rutuja Kshirsagar, Prasanth Shyamsundar, Gabriel Perdue
2022
Saved 2026-04-21
The computational power of real-world quantum computers is limited by errors. When using quantum computers to perform algorithms which cannot be efficiently simulated classically, it is important to quantify the accuracy with which the computation has been performed. In this work we introduce a machine-learning-based technique to estimate the fidelity between the state produced by a noisy quantum circuit and the target state corresponding to ideal noise-free computation. Our machine learning model is trained in a supervised manner, using smaller or simpler circuits for which the fidelity can be estimated using other techniques like direct fidelity estimation and quantum state tomography. We demonstrate that the trained model can predict the fidelities of more complicated circuits for which such methods are infeasible.
JRRUIFID
webpage
Saved 2026-04-21
Why Xanadu’s hype-averse QML team is betting on the quantum Fourier transform to advance machine learning and better machine learning models.
87RKDS9E
blogPost
Alex Tabarrok
2026
Saved 2026-04-21
It’s well known that grade inflation has “degraded” the informational content of grades at many colleges. At Harvard, two-thirds of all undergraduate grades are now A’s—up from about a quarter two decades ago. In response, a Harvard faculty committee has proposed capping A grades at 20 percent of each class (plus a cushion for small […]
LE77V83K
webpage
Tony Kulesa
2023
Saved 2026-04-21
Tyler Cowen is an economist that has spotted top talent in fields ranging from biotech to literature, often years before insiders. How does he do it?
VUZVQQTY
webpage
Substack
Saved 2026-04-21
addiction is good, actually.
XEEEG4DZ
webpage
Substack
Saved 2026-04-21
What will happen in our second peasanthood
NAHF6KQH
book
American Prometheus: The Triumph and Tragedy of J. Robert Oppenheimer
Kai Bird, Martin Sherwin
Saved 2026-04-20
QZAF8EG4
webpage
Craig Mod
Saved 2026-04-20
Essay on the benefits of speedy software, and how it affects user perception of engineering quality and overall usability
9R5JZPPD
webpage
Substack
Saved 2026-04-19
that's the artist's vocation for ya
9WU9WEJE
webpage
Saved 2026-04-19
Spawn coding agents that run infinitely in the cloud. Powered by AI SDK, Gateway, Sandbox, and Workflow SDK.
TJ4WXXPD
preprint
Christos Louizos, Max Welling, Diederik P. Kingma
2018
Saved 2026-04-19
We propose a practical method for $L_0$ norm regularization for neural networks: pruning the network during training by encouraging weights to become exactly zero. Such regularization is interesting since (1) it can greatly speed up training and inference, and (2) it can improve generalization. AIC and BIC, well-known model selection criteria, are special cases of $L_0$ regularization. However, since the $L_0$ norm of weights is non-differentiable, we cannot incorporate it directly as a regularization term in the objective function. We propose a solution through the inclusion of a collection of non-negative stochastic gates, which collectively determine which weights to set to zero. We show that, somewhat surprisingly, for certain distributions over the gates, the expected $L_0$ norm of the resulting gated weights is differentiable with respect to the distribution parameters. We further propose the \emph{hard concrete} distribution for the gates, which is obtained by "stretching" a binary concrete distribution and then transforming its samples with a hard-sigmoid. The parameters of the distribution over the gates can then be jointly optimized with the original network parameters. As a result our method allows for straightforward and efficient learning of model structures with stochastic gradient descent and allows for conditional computation in a principled way. We perform various experiments to demonstrate the effectiveness of the resulting approach and regularizer.
KT5885AI
preprint
Jiacheng Liu, Xiaohan Zhao, Xinyi Shang, Zhiqiang Shen
2026
Saved 2026-04-18
Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study describes its comprehensive architecture by analyzing the publicly available TypeScript source code and further comparing it with OpenClaw, an independent open-source AI agent system that answers many of the same design questions from a different deployment context. Our analysis identifies five human values, philosophies, and needs that motivate the architecture (human decision authority, safety and security, reliable execution, capability amplification, and contextual adaptability) and traces them through thirteen design principles to specific implementation choices. The core of the system is a simple while-loop that calls the model, runs tools, and repeats. Most of the code, however, lives in the systems around this loop: a permission system with seven modes and an ML-based classifier, a five-layer compaction pipeline for context management, four extensibility mechanisms (MCP, plugins, skills, and hooks), a subagent delegation mechanism with worktree isolation, and append-oriented session storage. A comparison with OpenClaw, a multi-channel personal assistant gateway, shows that the same recurring design questions produce different architectural answers when the deployment context changes: from per-action safety classification to perimeter-level access control, from a single CLI loop to an embedded runtime within a gateway control plane, and from context-window extensions to gateway-wide capability registration. We finally identify six open design directions for future agent systems, grounded in recent empirical, architectural, and policy literature.
AKWQFZ2H
webpage
Ana-Piligrim—Updated on Thursday, April 9, 2026
Saved 2026-04-18
In a time of speed and uncertainty of the future we often talk about ways to improve our productivity as designers yet not often nurture and build upon wha...
9G5DL4HD
blogPost
Dan Meyer
2024
Saved 2026-04-18
The chatbots don't know when to start chatting with learners.
DLUVGDRE
webpage
Substack
Saved 2026-04-18
How a Strategic Pivot, a 20% Project, and 1,500 Hours of Focused Effort Led to My Dream Job in AGI Safety.
RIKN2HQ3
webpage
Substack
Saved 2026-04-18
18 Months of Strategic Job Hunting
45JXQFIY
webpage
Saved 2026-04-18
A deep dive on the future of energy
Z7SNIXQR
webpage
Substack
Saved 2026-04-17
Jasmine Sun on writing and doing everything to win
2Y2QST3E
blogPost
0xsmac
2025
Saved 2026-04-17
Hypergambling and the boredom of real compounding
H2IUIN85
forumPost
Thariq [@trq212]
2026
Saved 2026-04-17
3HNWL8X8
webpage
Saved 2026-04-17
Brian Lovin is a designer and software engineer living in San Francisco, currently designing AI products at Notion.
ZJLTVMCR
webpage
Saved 2026-04-17
IT97AGQW
blogPost
Contributor
2012
Saved 2026-04-17
Editor’s note: Paul Stamatiou is Co-founder of Picplum, a Y Combinator-backed photo printing service, where he obsesses over both design and development. He also co-founded Notifo (YC W10) and Skribit. Follow him on his blog, PaulStamatiou.com, and on Twitter: @Stammy. Reminisce with me for a bit. Do you remember the first time you got an Internet connection? Before your computer was always connected and when going online was a thing you had to plan. The joys of seeing new browsers like Phoenix emerge. Your excitement when you first experienced the Web with your new high-speed connection. It was a time when sites rarely had any JavaScript and DHTML was the buzzword of the year. Now it's hard to believe that Chrome is just a few years old.
J3A7B4GP
blogPost
Karina Nguyen
2026
Saved 2026-04-17
With deep gratitude to the collaborators, mentors, and friends who shaped how I think about the world
IYLFK65L
webpage
`
Substack
Saved 2026-04-17
With deep gratitude to the collaborators, mentors, and friends who shaped how I think about the world
N2RMT9ZB
webpage
Substack
Saved 2026-04-17
Try to keep an open mind as the world gets increasingly wild.
BVQMVMHA
webpage
Substack
Saved 2026-04-17
Trust me, I'm an alcoholic. I really am just doing this for fun.
HGXSZKGN
webpage
Substack
Saved 2026-04-17
How I cold email billionaires and get responses (CEOs of Uber, Groupon, Coursera)
MRLVVH5E
blogPost
Mishti Sharma
2020
Saved 2026-04-17
people, not pedigrees, were my ticket to success
HL5752TZ
webpage
Substack
Saved 2026-04-17
people, not pedigrees, were my ticket to success
GGPHSL2T
webpage
Substack
Saved 2026-04-16
On hacker mindset
DY9U5E92
webpage
Saved 2026-04-15
Y6ZXEBQV
book
Compilers: principles, techniques, & tools
Alfred V. Aho, Alfred V. Aho
2007 · Pearson/Addison Wesley
Saved 2026-04-15
BN75S3FD
journalArticle
The party’s AI: How China’s new AI systems are reshaping human rights
2025
Saved 2026-04-15
IN8Y3PKX
webpage
S. E. Gyges
2025
Saved 2026-04-15
The “Stochastic Parrot” Argument is Both Wrong and Actively Harmful
EJJJHHYL
webpage
ppadjin
2026
Saved 2026-04-14
Pavle Pađin
EQZGZIFU
preprint
Jake Ward, Paul Riechers, Adam Shai
2025
Saved 2026-04-14
Reasoning models leverage inference-time compute to significantly enhance the performance of language models on difficult logical tasks, and have become a dominating paradigm in frontier LLMs. Despite their wide adoption, the mechanisms underpinning the enhanced performance of these reasoning models are not well understood. In this work, we show that the majority of new capabilities in reasoning models can be elicited by small, single-rank changes to base model parameters, with many of these changes being interpretable. Specifically, we use a rank-1 LoRA to create a minimal parameter adapter for Qwen-2.5-32B-Instruct which recovers 73-90% of reasoning-benchmark performance compared to a full parameter finetune. We find that the activations of this LoRA are as interpretable as MLP neurons, and fire for reasoning-specific behaviors. Finally, we train a sparse autoencoder on the entire activation state of this LoRA and identify fine-grained and monosemantic features. Our findings highlight that reasoning performance can arise largely from minimal changes to base model parameters, and explore what these changes affect. More broadly, our work shows that parameter-efficient training methods can be used as a targeted lens for uncovering fundamental insights about language model behavior and dynamics.
MJSG64KI
preprint
Adam Shai, Loren Amdahl-Culleton, Casper L. Christensen, Henry R. Bigelow, Fernando E. Rosas, Alexander B. Boyd, Eric A. Alt, Kyle J. Ray et al.
2026
Saved 2026-04-14
Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize two representational hypotheses: (1) a representation in the product space of all factors, whose dimension grows exponentially with the number of parts, or (2) a factored representation in orthogonal subspaces, whose dimension grows linearly. The factored representation is lossless when factors are conditionally independent, but sacrifices predictive fidelity otherwise, creating a tradeoff between dimensional efficiency and accuracy. We derive precise predictions about the geometric structure of activations for each, including the number of subspaces, their dimensionality, and the arrangement of context embeddings within them. We test between these hypotheses on transformers trained on synthetic processes with known latent structure. Models learn factored representations when factors are conditionally independent, and continue to favor them early in training even when noise or hidden dependencies undermine conditional independence, reflecting an inductive bias toward factoring at the cost of fidelity. This provides a principled explanation for why transformers decompose the world into parts, and suggests that interpretable low dimensional structure may persist even in models trained on complex data.
QF4NDGE7
webpage
Saved 2026-04-14
Simplex is an AI safety research organization building a science of intelligence.
77GUIVNK
journalArticle
Sequence to Sequence Learning with Neural Networks
Ilya Sutskever, Oriol Vinyals, Quoc V Le
Saved 2026-04-13
Deep Neural Networks (DNNs) are powerful models that have achieved excellent performance on difficult learning tasks. Although DNNs work well whenever large labeled training sets are available, they cannot be used to map sequences to sequences. In this paper, we present a general end-to-end approach to sequence learning that makes minimal assumptions on the sequence structure. Our method uses a multilayered Long Short-Term Memory (LSTM) to map the input sequence to a vector of a fixed dimensionality, and then another deep LSTM to decode the target sequence from the vector. Our main result is that on an English to French translation task from the WMT-14 dataset, the translations produced by the LSTM achieve a BLEU score of 34.8 on the entire test set, where the LSTM’s BLEU score was penalized on out-of-vocabulary words. Additionally, the LSTM did not have difficulty on long sentences. For comparison, a phrase-based SMT system achieves a BLEU score of 33.3 on the same dataset. When we used the LSTM to rerank the 1000 hypotheses produced by the aforementioned SMT system, its BLEU score increases to 36.5, which is close to the previous state of the art. The LSTM also learned sensible phrase and sentence representations that are sensitive to word order and are relatively invariant to the active and the passive voice. Finally, we found that reversing the order of the words in all source sentences (but not target sentences) improved the LSTM’s performance markedly, because doing so introduced many short term dependencies between the source and the target sentence which made the optimization problem easier.
RVKFYX84
preprint
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, Yunfeng Liu
2023
Saved 2026-04-13
Position encoding recently has shown effective in the transformer architecture. It enables valuable supervision for dependency modeling between elements at different positions of the sequence. In this paper, we first investigate various methods to integrate positional information into the learning process of transformer-based language models. Then, we propose a novel method named Rotary Position Embedding(RoPE) to effectively leverage the positional information. Specifically, the proposed RoPE encodes the absolute position with a rotation matrix and meanwhile incorporates the explicit relative position dependency in self-attention formulation. Notably, RoPE enables valuable properties, including the flexibility of sequence length, decaying inter-token dependency with increasing relative distances, and the capability of equipping the linear self-attention with relative position encoding. Finally, we evaluate the enhanced transformer with rotary position embedding, also called RoFormer, on various long text classification benchmark datasets. Our experiments show that it consistently overcomes its alternatives. Furthermore, we provide a theoretical analysis to explain some experimental results. RoFormer is already integrated into Huggingface: https://huggingface.co/docs/transformers/model_doc/roformer.
ZHG2GBYH
blogPost
Alex Iskold
2015
Saved 2026-04-13
I’ve been getting a large number of bad VC asks lately. What is a bad VC ask? A bad VC ask is an open-ended ask for an intro to a venture firm. Here are some examples: – Do you know any…
TTAI4MQU
webpage
Saved 2026-04-13
Getting an introduction is a basic thing that startup founders do pretty much every day.We talked about introductions in the post about business development tips , and also in this post about asking for introductions .One other thing we teach Techstars founders how to do properly is to send someth
ZQ48MKV4
preprint
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin
2023
Saved 2026-04-13
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.
MRP56VHP
webpage
2026
Saved 2026-04-12
In his classic Programming Perl — affectionately known to a generation of technologists as "the Camel Book" — Larry Wall famously wrote of the three virtues of a programmer as laziness, impatience, and hubris: If we’re going to talk about good software design, we have to talk about Laziness, Impatience, and Hubris, the basis of good software design. We’ve all fallen into the trap of using cut-and-paste when we should have defined a higher-level abstraction, if only just a loop or subroutine. To be sure, some folks have gone to the opposite extreme of defining ever-growing mounds of higher level abstractions when they should have used cut-and-paste. Generally, though, most of us need to think about using more abstraction rather than less.
6L7G3KEB
blogPost
Alexandra Phelan
2024
Saved 2026-04-12
August 2023
JYJQTM72
journalArticle
Angela Rosy Morgillo, Stefano Mangini, Marco Piastra, Chiara Macchiavello
2024 · Quantum Machine Intelligence
Saved 2026-04-12
Quantum noise is currently limiting efficient quantum information processing and computation. In this work, we consider the tasks of reconstructing and classifying quantum states corrupted by the action of an unknown noisy channel using classical feedforward neural networks. By framing reconstruction as a regression problem, we show how such an approach can be used to recover with fidelities exceeding 99% the noiseless density matrices of quantum states of up to three qubits undergoing noisy evolution, and we test its performance with both single-qubit (bit-flip, phase-flip, depolarising, and amplitude damping) and two-qubit quantum channels (correlated amplitude damping). Moreover, we also consider the task of distinguishing between different quantum noisy channels, and show how a neural network-based classifier is able to solve such a classification problem with perfect accuracy.
XUIMQE6A
journalArticle
Suguru Endo, Simon C. Benjamin, Ying Li
2018 · Physical Review X · American Physical Society
Saved 2026-04-12
It is vital to minimize the impact of errors for near-future quantum devices that will lack the resources for full fault tolerance. Two quantum error mitigation (QEM) techniques have been introduced recently, namely, error extrapolation [Y. Li and S. C. Benjamin, Phys. Rev. X 7, 021050 (2017); K. Temme et al., Phys. Rev. Lett. 119, 180509 (2017)] and quasiprobability decomposition [K. Temme et al., Phys. Rev. Lett. 119, 180509 (2017)]. To enable practical implementation of these ideas, here we account for the inevitable imperfections in the experimentalist’s knowledge of the error model itself. We describe a protocol for systematically measuring the effect of errors so as to design efficient QEM circuits. We find that the effect of localized Markovian errors can be fully eliminated by inserting or replacing some gates with certain single-qubit Clifford gates and measurements. Finally, having introduced an exponential variant of the extrapolation method we contrast the QEM techniques using exact numerical simulation of up to 19 qubits in the context of a “swap” test circuit. Our optimized methods dramatically reduce the circuit’s output error without increasing the qubit count.
EDJLWTVC
journalArticle
H. Chen, L. Wossnig, S. Severini, H. Neven, M. Mohseni
2020 · Quantum Machine Intelligence
Saved 2026-04-12
Recent results have demonstrated the successful applications of quantum-classical hybrid methods to train quantum circuits for a variety of machine learning tasks. A natural question to ask is consequentially whether we can also train such quantum circuits to discriminate quantum data, i.e., perform classification on data stored in form of quantum states. Although quantum mechanics fundamentally forbids deterministic discrimination of non-orthogonal states, we show in this work that it is possible to train a quantum circuit to discriminate such data with a trade-off between minimizing error rates and inconclusiveness rates of the classification tasks. Our approach achieves at the same time a performance which is close to the theoretically optimal values and a generalization ability to previously unseen quantum data. This generalization power hence distinguishes our work from previous circuit optimization results and furthermore provides an example of a quantum machine learning task that has inherently no classical analogue.
4Z4IU66C
journalArticle
Piotr Czarnik, Andrew Arrasmith, Patrick J. Coles, Lukasz Cincio
2021 · Quantum
Saved 2026-04-12
Achieving near-term quantum advantage will require accurate estimation of quantum observables despite significant hardware noise. For this purpose, we propose a novel, scalable error-mitigation method that applies to gate-based quantum computers. The method generates training data $\{X_i^{\text{noisy}},X_i^{\text{exact}}\}$ via quantum circuits composed largely of Clifford gates, which can be efficiently simulated classically, where $X_i^{\text{noisy}}$ and $X_i^{\text{exact}}$ are noisy and noiseless observables respectively. Fitting a linear ansatz to this data then allows for the prediction of noise-free observables for arbitrary circuits. We analyze the performance of our method versus the number of qubits, circuit depth, and number of non-Clifford gates. We obtain an order-of-magnitude error reduction for a ground-state energy problem on 16 qubits in an IBMQ quantum computer and on a 64-qubit noisy simulator.
CGJEZRX4
preprint
Karan Kendre
2025
Saved 2026-04-12
Quantum noise fundamentally limits the utility of near-term quantum devices, making error mitigation essential for practical quantum computation. While traditional quantum error correction codes require substantial qubit overhead and complex syndrome decoding, we propose a machine learning approach that directly reconstructs clean quantum states from noisy density matrices without additional qubits. We formulate quantum noise reduction as a supervised learning problem using a convolutional neural network (CNN) autoencoder architecture with a novel fidelity-aware composite loss function. Our method is trained and evaluated on a comprehensive synthetic dataset of 10,000 density matrices derived from random 5-qubit quantum circuits, encompassing five noise types (depolarizing, amplitude damping, phase damping, bit-flip, and mixed noise) across four intensity levels (0.05-0.20). The CNN successfully reconstructs quantum states across all noise conditions, achieving an average fidelity improvement from 0.298 to 0.774 (Δ = 0.476). Notably, the model demonstrates superior performance on complex mixed noise scenarios and higher noise intensities, with mixed noise showing the highest corrected fidelity (0.807) and improvement (0.567). The approach effectively preserves both diagonal elements (populations) and off-diagonal elements (quantum coherences), making it suitable for entanglement-dependent quantum algorithms. While phase damping presents fundamental information-theoretic limitations, our results suggest that CNN-based density matrix reconstruction offers a promising, resource-efficient alternative to traditional quantum error correction for NISQ-era devices. This data-driven approach could enable practical quantum advantage with fewer physical qubits than conventional error correction schemes require.
LRURVMFI
webpage
Annette Vee
2024
Saved 2026-04-12
Artificial intelligence ‘bots’ may be widely available to teach in the coming years. But will they be effective? An expert on technology in education weighs in.
7HM6DSXV
webpage
Saved 2026-04-12
UYTH7FVL
blogPost
Saved 2026-04-11
In Surrounded by Idiots author Thomas Erikson describes the famous DISC profiles, a methods to sort out the differences in human communication.
BW5D87QL
forumPost
Adrianna Lakatos [@adriannalakatos]
2026
Saved 2026-04-11
99LPCTBA
webpage
Saved 2026-04-11
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
SLP3K4Q3
forumPost
alex zhang [@a1zhang]
2026
Saved 2026-04-11
JAZJA5DM
webpage
2026
Saved 2026-04-11
EN8T8C6S
preprint
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang et al.
2024
Saved 2026-04-11
Mathematical reasoning poses a significant challenge for language models due to its complex and structured nature. In this paper, we introduce DeepSeekMath 7B, which continues pretraining DeepSeek-Coder-Base-v1.5 7B with 120B math-related tokens sourced from Common Crawl, together with natural language and code data. DeepSeekMath 7B has achieved an impressive score of 51.7% on the competition-level MATH benchmark without relying on external toolkits and voting techniques, approaching the performance level of Gemini-Ultra and GPT-4. Self-consistency over 64 samples from DeepSeekMath 7B achieves 60.9% on MATH. The mathematical reasoning capability of DeepSeekMath is attributed to two key factors: First, we harness the significant potential of publicly available web data through a meticulously engineered data selection pipeline. Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO.
4KY5HX43
preprint
Johannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento, Alexander Mordvintsev, Andrey Zhmoginov, Max Vladymyrov
2023
Saved 2026-04-11
At present, the mechanisms of in-context learning in Transformers are not well understood and remain mostly an intuition. In this paper, we suggest that training Transformers on auto-regressive objectives is closely related to gradient-based meta-learning formulations. We start by providing a simple weight construction that shows the equivalence of data transformations induced by 1) a single linear self-attention layer and by 2) gradient-descent (GD) on a regression loss. Motivated by that construction, we show empirically that when training self-attention-only Transformers on simple regression tasks either the models learned by GD and Transformers show great similarity or, remarkably, the weights found by optimization match the construction. Thus we show how trained Transformers become mesa-optimizers i.e. learn models by gradient descent in their forward pass. This allows us, at least in the domain of regression problems, to mechanistically understand the inner workings of in-context learning in optimized Transformers. Building on this insight, we furthermore identify how Transformers surpass the performance of plain gradient descent by learning an iterative curvature correction and learn linear models on deep data representations to solve non-linear regression tasks. Finally, we discuss intriguing parallels to a mechanism identified to be crucial for in-context learning termed induction-head (Olsson et al., 2022) and show how it could be understood as a specific case of in-context learning by gradient descent learning within Transformers. Code to reproduce the experiments can be found at https://github.com/google-research/self-organising-systems/tree/master/transformers_learn_icl_by_gd .
A6KSITYV
webpage
Saved 2026-04-11
WW6JQXCR
videoRecording
Alpine Investors
2026
Saved 2026-04-10
Graham shares the four truths that Alpine is using to guide its decisions as we navigate this tech revolution. This was speech was recorded at Alpine's 2025’s Growth Summit.
FQCXJQTD
book
Elon Musk
Walter Isaacson
Saved 2026-04-10
VD54QJV3
journalArticle
Sarah Elven, Jorge Luis Castañeda Núñez, Samantha de Martino, Michelle Dugas, Sayan Kundu
2025 · npj Climate Action · Nature Publishing Group
Saved 2026-04-10
De facto exclusion of vulnerable populations from markets for energy-efficient technologies can result in multiple barriers to access. For example, exclusion can lead to limited knowledge about available products, an inability to distinguish high-quality from low-quality devices, and limited options for financing, making products seem unobtainable. However, behaviorally informed interventions can offer promising solutions in such contexts, even where exclusion is the result of structural causes. This paper uses a randomized control trial to consider the potential of such interventions for refugees in Uganda in the context of certified solar markets. We evaluate a behaviorally-informed information and savings session embedded in Village Savings and Lending Association (VSLA) meetings, finding evidence for increased pursuit of certified solar products in the treatment group two months later. Results manifest through the barriers described, with increased knowledge, trust in solar companies, financial inclusion through savings group support, and aspirations mediating effects.
KQEE4QWC
webpage
Saved 2026-04-10
💡Putting yourself out there increases the surface area for potential opportunities
MHG9TQNJ
webpage
Saved 2026-04-10
"The amount of serendipity that will occur in your life is directly proportional to the degree to which you do something you're passionate about combined with the total number of people to whom this is effectively communicated."
DIW338H2
webpage
2020
Saved 2026-04-10
An illustrated therapy session between Humanity, Capitalism, and Post-Capitalism.
Z38SF2YE
webpage
2019
Saved 2026-04-09
We are still in the "touching the elephant" phase of the "ideological mindsets" that emerge from "The Transition" (away from late-stage-capitalism at the birth of the internet). In other words, we don't yet know what the most coherent and powerful mindsets are. https://twitter.com/iiterature/status/1048301983109668864We can view
9DCEC3Q6
webpage
Saved 2026-04-09
In-depth report on Building the Wisdom Age and creating the Wisdom Economy.
BFEJG939
webpage
Saved 2026-04-09
F98HF93W
webpage
Saved 2026-04-09
1/ Maps are one of the most effective ways to communicate Here's a 🧵of maps, starting with everyone's favorite — memespace maps https://xkcd.com/802/ https://www.ribbonfarm.com/you-are-here/ https://slatestarcodex.com/2018/07/19/sentimental-cartography/ https://www.lesswrong.com/posts/WzPJRNYWhMXQTEj69/a-map-of-bay-area-memespace
TZJK9EIQ
webpage
Nadia Asparouhova
2022
Saved 2026-04-09
Tech as a system of values, and not just an industry, is heavily driven by its subcultures and their ideologies. Where do these ideologies come from, and how do they influence what’s accomplished?
XQB32ERQ
webpage
2023
Saved 2026-04-09
Almost everything in life must be done with others. If you learn to relate well to others, you increase your professional leverage and the emotional health of your relationships. I. Understanding Others The first step to relating well to others is understanding them. The first step to understanding others is
FJ4U7UY3
webpage
2022
Saved 2026-04-09
1. Stop Comparing Yourself Against Old Models You might be unconsciously following the path of your parents like a doctor or lawyer. Don't. Optimize for rate-of-learning [https://rooteco.notion.site/What-Roote-Hires-For-f21facf226e248e6af61551ef182c933] and consciously reject the paths given to you. Patrick Collison gives this advice [https://patrickcollison.com/advice]: > If you're
4UEQNYRE
blogPost
Saved 2026-04-09
The philosopher Kwame Appiah writes that “in life, the challenge is not so much to figure out how best to play the game; the challenge is to figure out what game you’re playing.” When I try to figure out what game I’m playing, I see that for the last 25 years I have been playing […]
MQ27FYRU
webpage
Saved 2026-04-09
I’ve observed thousands of founders and thought a lot about what it takes to make a huge amount of money or to create something important. Usually, people start off wanting the former and end up...
XGPDKJSP
webpage
2023
Saved 2026-04-09
One of the most powerful ways to build agency is to have "almost too much self-belief." But this can go badly quickly: look at how SBF and Elizabeth Holmes built reality distortion fields around themselves, defrauded everyone, and ended up in jail. Some reality distortion is good—we want to
44ZCIN9Q
webpage
2024
Saved 2026-04-09
HIU8X9MT
webpage
Saved 2026-04-09
My answer:
9EZ5GJR4
webpage
2022
Saved 2026-04-09
In summary: * Write for a specific reader. * Write concisely. * When editing, look to cut. * Write in chunks. * Have fun with it. * Leverage the computer with tools like Grammarly, hotkeys, and ergonomic keyboards. 1. Alex Danco's 5 Writing Tips Write in parallel tracks using "meanwhile". Meanwhile establishes parallel tracks of thought.
5C279IXD
webpage
2022
Saved 2026-04-09
For the first 30 years of my life, I coasted and procrastinated often. I am pretty ashamed of it. 😔 In the last few months, I've beaten procrastination. I'm proud of this! ❤️ Here's how I did it. I hope it helps you. 1. Procrastination Is A Harmful Positive Feedback Loop Procrastination
YFX5AJA8
webpage
2022
Saved 2026-04-09
I've previously written a piece on Why [https://www.rhyslindmark.com/anki/] You Should Use Anki. [https://www.rhyslindmark.com/anki/] This piece is how. 1. Understand your why. Why do you want to remember anything? For me, it was reading lots of books but forgetting them. What's your goal?
EIP6XABY
webpage
2022
Saved 2026-04-09
If you're a knowledge worker doing Tyler Cowen calls an intellectual athlete training [https://marginalrevolution.com/marginalrevolution/2019/07/learn-like-an-athlete-knowledge-workers-should-train.html] , you should use Anki or some other form of spaced repetition. See the responses to Michael Nielsen's tweet here for all the other great folks that do this: > The
R6RTAXYG
webpage
2023
Saved 2026-04-09
1. First, remind yourself that the goal is not to simply to consume, but to contribute back to society. Build, write, teach, help others. That's why you learn in the first place. 2. Use Anki. Here's why. Here's how. 3. Join Twitter. Use these tips. Join these communities. Follow the
D2Y6BE2D
webpage
Packy McCormick
2025
Saved 2026-04-09
Why it's worth trying to make the world more optimistic
SM68U97J
journalArticle
Max Roser
2022 · Our World in Data
Saved 2026-04-09
It is wrong to think these three statements contradict each other. We need to see that they are all true to see that a better world is possible.
WWFCGNPH
blogPost
Hannah Ritchie
2022
Saved 2026-04-09
More than half of young people think "humanity is doomed" due to climate change. We need to reframe the narrative from doom and sacrifice, to one of opportunity.
KM29AS4Q
webpage
Simon Sarris
2025
Saved 2026-04-09
I have edited and expanded this in a newer post, you should read that instead: School is Not Enough
C2AFZ9NF
webpage
Saved 2026-04-09
74MYG275
webpage
2022
Saved 2026-04-09
Some examples of people with high stamina. Inspired by /fast. * Beyoncé: As a child: "When some of us wanted to go to the movies, when we did have our off time, she was in the studio sitting there by herself writing a record. I’ll never forget, were at an
HW5K3KHC
webpage
Saved 2026-04-09
Don't "work harder." Instead: Cut away distractions. Stop doing fake work. Ask for help. Plan. Solicit criticism. Automate what can be automated. Engage your mentors. Exercise. Get more sleep. Whatever you do, don't "work harder." It's pretty much never the answer.
3MQZG5KR
blogPost
Simon Sarris
2023
Saved 2026-04-09
GXKPET7V
webpage
2023
Saved 2026-04-09
A list of things you're allowed to do that you thought you weren't, or didn't even know you could.
UT9DRF32
webpage
Brie Wolfson
2022
Saved 2026-04-09
Nostalgia for another way of working
RSTYGPSF
blogPost
Tyler Cowen
2018
Saved 2026-04-09
Yesterday I had lunch with a former Ph.D student of mine, who is now highly successful and tenured at a very good school.  I was reminded that, over twenty years ago, I was Graduate Director of Admissions.  One of my favorite strategies was to take strong candidates who applied for Masters and also offer them […]
STUSPPHN
webpage
2020
Saved 2026-04-09
YDJBC3PY
blogPost
Nabeel S. Qureshi
2025
Saved 2026-04-09
// This was originally a Google Doc where I gathered hard-won life lessons; eventually I open sourced it. It got a great response, including a shoutout from Tim Ferriss, so here it is on Substack.
JGHCC46W
webpage
Saved 2026-04-09
@_TamaraWinter: Recently I’ve gotten a lot of inbound from new grads asking for career advice, so I wrote down a list of things I wish I’d internalized sooner. The key thing: being precocious has an expiration...…
3BD7SFWA
blogPost
Rhys Lindmark
2023
Saved 2026-04-09
I give you permission. Permission to be ambitious. Permission to lead. Permission to apply for that fellowship, drop out of school, and start that startup.
ZQBLFPF9
webpage
2023
Saved 2026-04-08
Education is the kindling of a flame, not the filling of a vessel. — Socrates Compounding is everything. Compound interest dominates financial markets, compounding network effects dominate business, and compounding yourself dominates all other self-work you can do. So how do you compound yourself? We know what compounding looks like—an
NWBE5IRZ
webpage
Saved 2026-04-08
An interesting note: I asked many people I respect to also share their advice, and most of it ended up being completely different from what both I wrote and within the group of respondents. So, I...
LQ682VGF
webpage
Saved 2026-04-08
K56F3UNG
report
Claude Mythos Preview System Card
Saved 2026-04-08
5B5HYP8T
webpage
Saved 2026-04-08
N4CT8837
journalArticle
Heat, Light and Power for Refugees
Glada Lahn, Owen Grafham, Foreword Kofi Annan
· Saving Lives
Saved 2026-04-08
FMH6RCQB
webpage
2016
Saved 2026-04-07
Ideas for getting into compilers and programming languages
AWHD7A84
preprint
Varun Singh, Lucas Krauss, Sami Jaghouar, Matej Sirovatka, Charles Goddard, Fares Obied, Jack Min Ong, Jannik Straube et al.
2026
Saved 2026-04-07
We present the technical report for Arcee Trinity Large, a sparse Mixture-of-Experts model with 400B total parameters and 13B activated per token. Additionally, we report on Trinity Nano and Trinity Mini, with Trinity Nano having 6B total parameters with 1B activated per token, Trinity Mini having 26B total parameters with 3B activated per token. The models' modern architecture includes interleaved local and global attention, gated attention, depth-scaled sandwich norm, and sigmoid routing for Mixture-of-Experts. For Trinity Large, we also introduce a new MoE load balancing strategy titled Soft-clamped Momentum Expert Bias Updates (SMEBU). We train the models using the Muon optimizer. All three models completed training with zero loss spikes. Trinity Nano and Trinity Mini were pre-trained on 10 trillion tokens, and Trinity Large was pre-trained on 17 trillion tokens. The model checkpoints are available at https://huggingface.co/arcee-ai.
BPT8GJ7H
journalArticle
Stephanie Hirmer, Heather Cruickshank
2014 · Renewable and Sustainable Energy Reviews
Saved 2026-04-07
User-value is a determining factor for product acceptance in product design. Research on rural electrification to date, however, does not draw sufficient attention to the importance of user-value with regard to the overall success of a project. This is evident from the analysis of project reports and applicable indicators from agencies active in the sector. Learning from the design, psychology and sociology literatures, it is important that rural electrification projects incorporate the value perception of the end-user and extend their success beyond the commonly used criteria of financial value, the appropriateness of the technology, capacity building and technology uptake. Creating value for the enduser is particularly important for project acceptance and the sustainability of a scheme once it has been handed over to the local community. In this research paper, existing theories and models of value-theory are transposed and applied to community operated rural electrification schemes and a user-value framework is developed. Furthermore, the importance of value to the end-user is clarified. Current literature on product design reveals that user-value has different properties, many of which are applicable to rural electrification. Five value pillars and their sub-categories important for the users of rural electrification projects are identified, namely: functional; social significance; epistemic; emotional; and cultural values. These pillars provide the main structure for the conceptual framework developed in this research paper. It is proposed that by targeting the values of the end-user, the key factors of uservalue applicable to rural electrification projects will be identified and the sustainability of the project will be better ensured.
973ZT33N
preprint
Marc Finzi, Shikai Qiu, Yiding Jiang, Pavel Izmailov, J. Zico Kolter, Andrew Gordon Wilson
2026
Saved 2026-04-06
Can we learn more from data than existed in the generating process itself? Can new and useful information be constructed from merely applying deterministic transformations to existing data? Can the learnable content in data be evaluated without considering a downstream task? On these questions, Shannon information and Kolmogorov complexity come up nearly empty-handed, in part because they assume observers with unlimited computational capacity and do not target the useful information content. In this work, we identify and exemplify three seeming paradoxes in information theory: (1) information cannot be increased by deterministic transformations; (2) information is independent of the order of data; (3) likelihood modeling is merely distribution matching. To shed light on the tension between these results and modern practice, and to quantify the value of data, we introduce epiplexity, a formalization of information capturing what computationally bounded observers can learn from data. Epiplexity captures the structural content in data while excluding time-bounded entropy, the random unpredictable content exemplified by pseudorandom number generators and chaotic dynamical systems. With these concepts, we demonstrate how information can be created with computation, how it depends on the ordering of the data, and how likelihood modeling can produce more complex programs than present in the data generating process itself. We also present practical procedures to estimate epiplexity which we show capture differences across data sources, track with downstream performance, and highlight dataset interventions that improve out-of-distribution generalization. In contrast to principles of model selection, epiplexity provides a theoretical foundation for data selection, guiding how to select, generate, or transform data for learning systems.
4NC253YK
webpage
Saved 2026-04-06
View recent discussion. Abstract: Can we learn more from data than existed in the generating process itself? Can new and useful information be constructed from merely applying deterministic transformations to existing data? Can the learnable content in data be evaluated without considering a downstream task? On these questions, Shannon information and Kolmogorov complexity come up nearly empty-handed, in part because they assume observers with unlimited computational capacity and do not target the useful information content. In this work, we identify and exemplify three seeming paradoxes in information theory: (1) information cannot be increased by deterministic transformations; (2) information is independent of the order of data; (3) likelihood modeling is merely distribution matching. To shed light on the tension between these results and modern practice, and to quantify the value of data, we introduce epiplexity, a formalization of information capturing what computationally bounded observers can learn from data. Epiplexity captures the structural content in data while excluding time-bounded entropy, the random unpredictable content exemplified by pseudorandom number generators and chaotic dynamical systems. With these concepts, we demonstrate how information can be created with computation, how it depends on the ordering of the data, and how likelihood modeling can produce more complex programs than present in the data generating process itself. We also present practical procedures to estimate epiplexity which we show capture differences across data sources, track with downstream performance, and highlight dataset interventions that improve out-of-distribution generalization. In contrast to principles of model selection, epiplexity provides a theoretical foundation for data selection, guiding how to select, generate, or transform data for learning systems.
85F6AUR7
webpage
Yuxin Wu
2025
Saved 2026-04-06
In Python, when a set of objects constructs a reference cycle, none of them would reach a zero refcount. In this case, even if these objects all go out-of-scope and are no longer accessible, they will
6W4J2UJI
preprint
Keiron O'Shea, Ryan Nash
2015
Saved 2026-04-06
The field of machine learning has taken a dramatic twist in recent times, with the rise of the Artificial Neural Network (ANN). These biologically inspired computational models are able to far exceed the performance of previous forms of artificial intelligence in common machine learning tasks. One of the most impressive forms of ANN architecture is that of the Convolutional Neural Network (CNN). CNNs are primarily used to solve difficult image-driven pattern recognition tasks and with their precise yet simple architecture, offers a simplified method of getting started with ANNs. This document provides a brief introduction to CNNs, discussing recently published papers and newly formed techniques in developing these brilliantly fantastic image recognition models. This introduction assumes you are familiar with the fundamentals of ANNs and machine learning.
NG3SXSJL
preprint
Wilson Yan, Yunzhi Zhang, Pieter Abbeel, Aravind Srinivas
2021
Saved 2026-04-06
We present VideoGPT: a conceptually simple architecture for scaling likelihood based generative modeling to natural videos. VideoGPT uses VQVAE that learns downsampled discrete latent representations of a raw video by employing 3D convolutions and axial self-attention. A simple GPTlike architecture is then used to autoregressively model the discrete latents using spatio-temporal position encodings. Despite the simplicity in formulation and ease of training, our architecture is able to generate samples competitive with stateof-the-art GAN models for video generation on the BAIR Robot dataset, and generate high fidelity natural videos from UCF-101 and Tumbler GIF Dataset (TGIF). We hope our proposed architecture serves as a reproducible reference for a minimalistic implementation of transformer based video generation models. Samples and code are available at https://wilson1yan. github.io/videogpt/index.html.
Y3DNPI89
blogPost
Brian Mattson
2025
Saved 2026-04-05
No.262: Friday, June 13, 2025
FH6K9XJ5
webpage
kmathison6
2025
Saved 2026-04-05
Throughout its history, the Westminster Theological Journal published many articles written by Cornelius Van Til, and it also published many articles written about Van Til. The Spring 1995 issue of the journal, which commemorated the one hundredth birthday of Van Til, was unique in that the entire issue was devoted to articles related to his system of thought. This issue included articles by notable presuppositionalist scholars such as Greg L. Bahnsen, John M. Frame, William Edgar, K. Scott Olip
7JH5CHIK
blogPost
Simons Institute Editor
2024
Saved 2026-04-05
3B57KYPK
preprint
Dawei Zhu, Xiaoyu Shen, Marius Mosbach, Andreas Stephan, Dietrich Klakow
2023
Saved 2026-04-04
Weakly supervised learning is a popular approach for training machine learning models in low-resource settings. Instead of requesting high-quality yet costly human annotations, it allows training models with noisy annotations obtained from various weak sources. Recently, many sophisticated approaches have been proposed for robust training under label noise, reporting impressive results. In this paper, we revisit the setup of these approaches and find that the benefits brought by these approaches are significantly overestimated. Specifically, we find that the success of existing weakly supervised learning approaches heavily relies on the availability of clean validation samples which, as we show, can be leveraged much more efficiently by simply training on them. After using these clean labels in training, the advantages of using these sophisticated approaches are mostly wiped out. This remains true even when reducing the size of the available clean data to just five samples per class, making these approaches impractical. To understand the true value of weakly supervised learning, we thoroughly analyze diverse NLP datasets and tasks to ascertain when and why weakly supervised approaches work. Based on our findings, we provide recommendations for future research.
65M4BJDL
preprint
Rishi Jha, Collin Zhang, Vitaly Shmatikov, John X. Morris
2026
Saved 2026-04-04
We introduce the first method for translating text embeddings from one vector space to another without any paired data, encoders, or predefined sets of matches. Our unsupervised approach translates any embedding to and from a universal latent representation (i.e., a universal semantic structure conjectured by the Platonic Representation Hypothesis). Our translations achieve high cosine similarity across model pairs with different architectures, parameter counts, and training datasets. The ability to translate unknown embeddings into a different space while preserving their geometry has serious implications for the security of vector databases. An adversary with access only to embedding vectors can extract sensitive information about the underlying documents, sufficient for classification and attribute inference.
HXNZ6I8Y
preprint
Minyoung Huh, Brian Cheung, Tongzhou Wang, Phillip Isola
2024
Saved 2026-04-04
We argue that representations in AI models, particularly deep networks, are converging. First, we survey many examples of convergence in the literature: over time and across multiple domains, the ways by which different neural networks represent data are becoming more aligned. Next, we demonstrate convergence across data modalities: as vision models and language models get larger, they measure distance between datapoints in a more and more alike way. We hypothesize that this convergence is driving toward a shared statistical model of reality, akin to Plato's concept of an ideal reality. We term such a representation the platonic representation and discuss several possible selective pressures toward it. Finally, we discuss the implications of these trends, their limitations, and counterexamples to our analysis.
UTEACSNT
preprint
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas Müller, Harry Saini, Yam Levi, Dominik Lorenz et al.
2024
Saved 2026-04-04
Diffusion models create data from noise by inverting the forward paths of data towards noise and have emerged as a powerful generative modeling technique for high-dimensional, perceptual data such as images and videos. Rectified flow is a recent generative model formulation that connects data and noise in a straight line. Despite its better theoretical properties and conceptual simplicity, it is not yet decisively established as standard practice. In this work, we improve existing noise sampling techniques for training rectified flow models by biasing them towards perceptually relevant scales. Through a large-scale study, we demonstrate the superior performance of this approach compared to established diffusion formulations for high-resolution text-to-image synthesis. Additionally, we present a novel transformer-based architecture for text-to-image generation that uses separate weights for the two modalities and enables a bidirectional flow of information between image and text tokens, improving text comprehension, typography, and human preference ratings. We demonstrate that this architecture follows predictable scaling trends and correlates lower validation loss to improved text-to-image synthesis as measured by various metrics and human evaluations. Our largest models outperform state-of-the-art models, and we will make our experimental data, code, and model weights publicly available.
EN8IJZQ6
webpage
Saved 2026-04-04
Powered by our Oasis video model, your actions shape the environment in real time.
IUXBH444
journalArticle
Decart, Julian Quevedo, Quinn McIntyre, Spruce Campbell, Robert Wachen
2024
Saved 2026-04-04
We're excited to announce Oasis, the first playable, realtime, open-world AI model. It's a video game, but entirely generated by AI. Oasis is the first step in our research towards more complex interactive worlds. Oasis takes in user keyboard input and generates real-time gameplay, including physics, game rules, and graphics. You can move around, jump, pick up items, break blocks, and more. There is no game engine; just a foundation model. We believe fast transformer inference is the missing link to making generative video a reality. Using Decart's inference engine, we show that real-time video is possible. When Etched's transformer ASIC, Sohu, is released, we can run models like Oasis in 4K. Today, we're releasing Oasis's code, the weights of a 500M parameter model you can run locally, and a live playable demo of a larger checkpoint.
5B2QM7FS
preprint
Mingxing Xu, Wenrui Dai, Chunmiao Liu, Xing Gao, Weiyao Lin, Guo-Jun Qi, Hongkai Xiong
2021
Saved 2026-04-03
Traffic forecasting has emerged as a core component of intelligent transportation systems. However, timely accurate traffic forecasting, especially long-term forecasting, still remains an open challenge due to the highly nonlinear and dynamic spatial-temporal dependencies of traffic flows. In this paper, we propose a novel paradigm of Spatial-Temporal Transformer Networks (STTNs) that leverages dynamical directed spatial dependencies and long-range temporal dependencies to improve the accuracy of long-term traffic forecasting. Specifically, we present a new variant of graph neural networks, named spatial transformer, by dynamically modeling directed spatial dependencies with self-attention mechanism to capture realtime traffic conditions as well as the directionality of traffic flows. Furthermore, different spatial dependency patterns can be jointly modeled with multi-heads attention mechanism to consider diverse relationships related to different factors (e.g. similarity, connectivity and covariance). On the other hand, the temporal transformer is utilized to model long-range bidirectional temporal dependencies across multiple time steps. Finally, they are composed as a block to jointly model the spatial-temporal dependencies for accurate traffic prediction. Compared to existing works, the proposed model enables fast and scalable training over a long range spatial-temporal dependencies. Experiment results demonstrate that the proposed model achieves competitive results compared with the state-of-the-arts, especially forecasting long-term traffic flows on real-world PeMS-Bay and PeMSD7(M) datasets.
GAW5DBN4
preprint
Tero Karras, Miika Aittala, Timo Aila, Samuli Laine
2022
Saved 2026-04-03
We argue that the theory and practice of diffusion-based generative models are currently unnecessarily convoluted and seek to remedy the situation by presenting a design space that clearly separates the concrete design choices. This lets us identify several changes to both the sampling and training processes, as well as preconditioning of the score networks. Together, our improvements yield new state-of-the-art FID of 1.79 for CIFAR-10 in a class-conditional setting and 1.97 in an unconditional setting, with much faster sampling (35 network evaluations per image) than prior designs. To further demonstrate their modular nature, we show that our design changes dramatically improve both the efficiency and quality obtainable with pre-trained score networks from previous work, including improving the FID of a previously trained ImageNet-64 model from 2.07 to near-SOTA 1.55, and after re-training with our proposed improvements to a new SOTA of 1.36.
7F2ZM26S
preprint
Luozhou Wang, Zhifei Chen, Yihua Du, Dongyu Yan, Wenhang Ge, Guibao Shen, Xinli Xu, Leyi Wu et al.
2026
Saved 2026-04-02
Large-scale video generation models have demonstrated emergent physical coherence, positioning them as potential world models. However, a gap remains between contemporary "stateless" video architectures and classic state-centric world model theories. This work bridges this gap by proposing a novel taxonomy centered on two pillars: State Construction and Dynamics Modeling. We categorize state construction into implicit paradigms (context management) and explicit paradigms (latent compression), while dynamics modeling is analyzed through knowledge integration and architectural reformulation. Furthermore, we advocate for a transition in evaluation from visual fidelity to functional benchmarks, testing physical persistence and causal reasoning. We conclude by identifying two critical frontiers: enhancing persistence via data-driven memory and compressed fidelity, and advancing causality through latent factor decoupling and reasoning-prior integration. By addressing these challenges, the field can evolve from generating visually plausible videos to building robust, general-purpose world simulators.
W78XERGW
blogPost
Bryan Johnson
2021
Saved 2026-04-01
Dear Reader,
SVJ7VRQ8
webpage
Saved 2026-04-01
F4U2CYBZ
webpage
Saved 2026-04-01
YUCT8HHH
webpage
Saved 2026-04-01
WWK79LEN
webpage
George Mack
Saved 2026-03-30
High agency might be the most important idea of the 21st century. This essay is what I wish I read at 18, rather than writing at 30. Join me in the high agency rabbit hole.
RH3ESCAV
webpage
Saved 2026-03-30
YS64ZVWA
webpage
Saved 2026-03-19
Lecture notes with an introduction to machine learning and discussion of linear classification and the perceptron update rule.
AB9LJDZ8
journalArticle
Bridging Usability and Performance: A Tensor Compiler for Autovectorizing Homomorphic Encryption
Edward Chen, Fraser Brown, Wenting Zheng
Saved 2026-03-19
Homomorphic encryption (HE) offers strong privacy guarantees by enabling computation over encrypted data. However, the performance of tensor operations in HE is highly sensitive to how the plaintext data is packed into ciphertexts. Large tensor programs introduce numerous possible layout assignments, making it both challenging and tedious for users to manually write efficient HE programs. In this paper, we present Rotom, a compilation framework that autovectorizes tensor programs into optimized HE programs. Rotom systematically explores a wide range of layout assignments, applies state-of-the-art optimizations, and automatically generates an equivalent, efficient HE program. At its core, Rotom utilizes a novel, lightweight ApplyRoll layout conversion operator to easily modify the underlying data layouts and unlock new avenues for performance gains. Our evaluation demonstrates Rotom scalably compiles all tensor workloads in under 5 minutes, reduces rotations in hand-tuned protocols by up to 3×, and achieves up to 80× performance improvement over prior autovectorization systems.
P4PTXZYP
preprint
Tomás Gonzalez, Giulia Fanti, Aaditya Ramdas
2025
Saved 2026-03-18
Private Evolution (PE) is a promising training-free method for differentially private (DP) synthetic data generation. While it achieves strong performance in some domains (e.g., images and text), its behavior in others (e.g., tabular data) is less consistent. To date, the only theoretical analysis of the convergence of PE depends on unrealistic assumptions about both the algorithm’s behavior and the structure of the sensitive dataset. In this work, we develop a new theoretical framework to understand PE’s practical behavior and identify sufficient conditions for its convergence. For d-dimensional sensitive datasets with n data points from a convex and compact domain, we prove that under the right hyperparameter settings and given access to the Gaussian variation API proposed in [33], PE produces an (ε, δ)-DP synthetic dataset with expected 1-Wasserstein distance O˜(d(nε)−1/d) from the original; this establishes worst-case convergence of the algorithm as n → ∞. Our analysis extends to general Banach spaces as well. We also connect PE to the Private Signed Measure Mechanism, a method for DP synthetic data generation that has thus far not seen much practical adoption. We demonstrate the practical relevance of our theoretical findings in experiments.
N6YJP5AC
preprint
Reza Shokri, Marco Stronati, Congzheng Song, Vitaly Shmatikov
2017
Saved 2026-03-13
We quantitatively investigate how machine learning models leak information about the individual data records on which they were trained. We focus on the basic membership inference attack: given a data record and black-box access to a model, determine if the record was in the model's training dataset. To perform membership inference against a target model, we make adversarial use of machine learning and train our own inference model to recognize differences in the target model's predictions on the inputs that it trained on versus the inputs that it did not train on. We empirically evaluate our inference techniques on classification models trained by commercial "machine learning as a service" providers such as Google and Amazon. Using realistic datasets and classification tasks, including a hospital discharge dataset whose membership is sensitive from the privacy perspective, we show that these models can be vulnerable to membership inference attacks. We then investigate the factors that influence this leakage and evaluate mitigation strategies.
L4LB8FNK
blogPost
Harshwardhan Fartale
2023
Saved 2026-03-13
A comprehensive overview of various libraries and frameworks for differential privacy and their use cases.
YMXA9J3I
webpage
Saved 2026-03-13
Posted by Brendan McMahan and Daniel Ramage, Research ScientistsStandard machine learning approaches require centralizing the training data on one ...
XDHLS69S
book
The Ethical Algorithm
Saved 2026-03-13
8R4FZVBC
book
Untitled
Saved 2026-03-13
GEQI6TAL
blogPost
development
2025
Saved 2026-03-12
In the book “Good Inside” by Dr. Becky Kennedy (2022), we were reminded of an important principle and tool that can help cut this circuit off early, and lead to happier and healthier relationships. Today on the blog we're talking about the Most Generous Interpretation.
WSTUIPPQ
preprint
Daniel Kang, Tatsunori Hashimoto, Ion Stoica, Yi Sun
2022
Saved 2026-03-12
As ML models have increased in capabilities and accuracy, so has the complexity of their deployments. Increasingly, ML model consumers are turning to service providers to serve the ML models in the ML-as-a-service (MLaaS) paradigm. As MLaaS proliferates, a critical requirement emerges: how can model consumers verify that the correct predictions were served, in the face of malicious, lazy, or buggy service providers? In this work, we present the first practical ImageNet-scale method to verify ML model inference non-interactively, i.e., after the inference has been done. To do so, we leverage recent developments in ZK-SNARKs (zero-knowledge succinct non-interactive argument of knowledge), a form of zero-knowledge proofs. ZK-SNARKs allows us to verify ML model execution non-interactively and with only standard cryptographic hardness assumptions. In particular, we provide the first ZK-SNARK proof of valid inference for a full resolution ImageNet model, achieving 79\% top-5 accuracy. We further use these ZK-SNARKs to design protocols to verify ML model execution in a variety of scenarios, including for verifying MLaaS predictions, verifying MLaaS model accuracy, and using ML models for trustless retrieval. Together, our results show that ZK-SNARKs have the promise to make verified ML model inference practical.
NUM9FHLR
webpage
Saved 2026-03-11
Most people will spend decades in chronic pain to avoid a few minutes of acute pain.
YFIFDRQM
webpage
Saved 2026-03-11
Some tips on how to get a lot done with little time.
DG8IAIG5
preprint
H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, Blaise Agüera y Arcas
2023
Saved 2026-03-10
Modern mobile devices have access to a wealth of data suitable for learning models, which in turn can greatly improve the user experience on the device. For example, language models can improve speech recognition and text entry, and image models can automatically select good photos. However, this rich data is often privacy sensitive, large in quantity, or both, which may preclude logging to the data center and training there using conventional approaches. We advocate an alternative that leaves the training data distributed on the mobile devices, and learns a shared model by aggregating locally-computed updates. We term this decentralized approach Federated Learning. We present a practical method for the federated learning of deep networks based on iterative model averaging, and conduct an extensive empirical evaluation, considering five different model architectures and four datasets. These experiments demonstrate the approach is robust to the unbalanced and non-IID data distributions that are a defining characteristic of this setting. Communication costs are the principal constraint, and we show a reduction in required communication rounds by 10-100x as compared to synchronized stochastic gradient descent.
URL8UBRZ
journalArticle
Cynthia Dwork, Aaron Roth
2014 · Foundations and Trends® in Theoretical Computer Science
Saved 2026-03-10
The problem of privacy-preserving data analysis has a long history spanning multiple disciplines. As electronic data about individuals becomes increasingly detailed, and as technology enables ever more powerful collection and curation of these data, the need increases for a robust, meaningful, and mathematically rigorous definition of privacy, together with a computationally rich class of algorithms that satisfy this definition. Differential Privacy is such a definition. After motivating and discussing the meaning of differential privacy, the preponderance of this monograph is devoted to fundamental techniques for achieving differential privacy, and application of these techniques in creative combinations, using the query-release problem as an ongoing example. A key point is that, by rethinking the computational goal, one can often obtain far better results than would be achieved by methodically replacing each step of a non-private computation with a differentially private implementation. Despite some astonishingly powerful computational results, there are still fundamental limitations — not just on what can be achieved with differential privacy but on what can be achieved with any method that protects against a complete breakdown in privacy. Virtually all the algorithms discussed herein maintain differential privacy against adversaries of arbitrary computational power. Certain algorithms are computationally intensive, others are efficient. Computational complexity for the adversary and the algorithm are both discussed. We then turn from fundamentals to applications other than queryrelease, discussing differentially private methods for mechanism design and machine learning. The vast majority of the literature on differentially private algorithms considers a single, static, database that is subject to many analyses. Differential privacy in other models, including distributed databases and computations on data streams is discussed. Finally, we note that this work is meant as a thorough introduction to the problems and techniques of differential privacy, but is not intended to be an exhaustive survey — there is by now a vast amount of work in differential privacy, and we can cover only a small portion of it.
Z2TQ65FD
webpage
Saved 2026-03-10
Because our personalities are malleable, the internet lets us become anyone.
DNN4UWM3
journalArticle
Scott Alexander
2012
Saved 2026-03-10
Slippery slopes are themselves a slippery concept. Imagine trying to explain them to an alien: …
UG2RME2J
webpage
Saved 2026-03-10
Looking foolish is underrated.
XZUJ2L9A
book
Test-driven development: by example
Kent Beck
2015 · Addison-Wesley
Saved 2026-03-05
WHBRA9HP
journalArticle
K. De Raedt, K. Michielsen, H. De Raedt, B. Trieu, G. Arnold, M. Richter, Th Lippert, H. Watanabe et al.
2007 · Computer Physics Communications
Saved 2026-03-05
We describe portable software to simulate universal quantum computers on massive parallel computers. We illustrate the use of the simulation software by running various quantum algorithms on different computer architectures, such as a IBM BlueGene/L, a IBM Regatta p690+, a Hitachi SR11000/J1, a Cray X1E, a SGI Altix 3700 and clusters of PCs running Windows XP. We study the performance of the software by simulating quantum computers containing up to 36 qubits, using up to 4096 processors and up to 1 TB of memory. Our results demonstrate that the simulator exhibits nearly ideal scaling as a function of the number of processors and suggest that the simulation software described in this paper may also serve as benchmark for testing high-end parallel computers.
45SAQXLC
journalArticle
High-Performance Interconnect-Aware MPI communication for Deep Learning Workloads
Yıltan Hassan Temucin
Saved 2026-03-05
KBFUZX44
preprint
Tyson Jones, Bálint Koczor, Simon C. Benjamin
2023
Saved 2026-03-05
Classical simulation of quantum computers is an irreplaceable step in the design of quantum algorithms. Exponential simulation costs demand the use of high-performance computing techniques, and in particular distribution, whereby the quantum state description is partitioned between a network of cooperating computers - necessary for the exact simulation of more than approximately 30 qubits. Distributed computing is notoriously difficult, requiring bespoke algorithms dissimilar to their serial counterparts with different resource considerations, and which appear to restrict the utilities of a quantum simulator. This manuscript presents a plethora of novel algorithms for distributed full-state simulation of gates, operators, noise channels and other calculations in digital quantum computers. We show how a simple, common but seemingly restrictive distribution model actually permits a rich set of advanced facilities including Pauli gadgets, many-controlled many-target general unitaries, density matrices, general decoherence channels, and partial traces. These algorithms include asymptotically, polynomially improved simulations of exotic gates, and thorough motivations for high-performance computing techniques which will be useful for even non-distributed simulators. Our results are derived in language familiar to a quantum information theory audience, and our algorithms formalised for the scientific simulation community. We have implemented all algorithms herein presented into an isolated, minimalist C++ project, hosted open-source on Github with a permissive MIT license, and extensive testing. This manuscript aims both to significantly improve the high-performance quantum simulation tools available, and offer a thorough introduction to, and derivation of, full-state simulation techniques.
KZ3DFQRH
book
Designing Data-Intensive Applications: The Big Ideas Behind Reliable, Scalable, and Maintainable Systems
Martin Kleppmann
2017 · "O'Reilly Media, Inc."
Saved 2026-03-04
Data is at the center of many challenges in system design today. Difficult issues need to be figured out, such as scalability, consistency, reliability, efficiency, and maintainability. In addition, we have an overwhelming variety of tools, including relational databases, NoSQL datastores, stream or batch processors, and message brokers. What are the right choices for your application? How do you make sense of all these buzzwords? In this practical and comprehensive guide, author Martin Kleppmann helps you navigate this diverse landscape by examining the pros and cons of various technologies for processing and storing data. Software keeps changing, but the fundamental principles remain the same. With this book, software engineers and architects will learn how to apply those ideas in practice, and how to make full use of data in modern applications. Peer under the hood of the systems you already use, and learn how to use and operate them more effectively Make informed decisions by identifying the strengths and weaknesses of different tools Navigate the trade-offs around consistency, scalability, fault tolerance, and complexity Understand the distributed systems research upon which modern databases are built Peek behind the scenes of major online services, and learn from their architectures
PXFQPZG6
journalArticle
CUDA C++ Programming Guide
Saved 2026-02-27
PPPGDKGL
blogPost
Dwmpl_Web_Hostinger24
2024
Saved 2026-02-24
E-waste contaminates soil and water, threatening agriculture and human health. Learn how proper recycling and stricter regulations can mitigate its impact.
GJ3I2QQD
journalArticle
Zi-Yu Huang, Chia-Chin Chiang, Jian-Hao Chen, Yi-Chian Chen, Hsin-Lung Chung, Yu-Ping Cai, Hsiu-Chuan Hsu
2023 · Scientific Reports · Nature Publishing Group
Saved 2026-02-24
Artificial intelligence has been successfully applied in various fields, one of which is computer vision. In this study, a deep neural network (DNN) was adopted for Facial emotion recognition (FER). One of the objectives in this study is to identify the critical facial features on which the DNN model focuses for FER. In particular, we utilized a convolutional neural network (CNN), the combination of squeeze-and-excitation network and the residual neural network, for the task of FER. We utilized AffectNet and the Real-World Affective Faces Database (RAF-DB) as the facial expression databases that provide learning samples for the CNN. The feature maps were extracted from the residual blocks for further analysis. Our analysis shows that the features around the nose and mouth are critical facial landmarks for the neural networks. Cross-database validations were conducted between the databases. The network model trained on AffectNet achieved 77.37% accuracy when validated on the RAF-DB, while the network model pretrained on AffectNet and then transfer learned on the RAF-DB results in validation accuracy of 83.37%. The outcomes of this study would improve the understanding of neural networks and assist with improving computer vision accuracy.
CWTUATB9
book
Programming Massively Parallel Processors
David Kirk, Hwu Wen-mei
Saved 2026-02-20
2DTM4YZ8
preprint
Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab Mirrokni
2025
Saved 2026-02-19
Despite the recent progresses, particularly in developing Language Models, there are fundamental challenges and unanswered questions about how such models can continually learn/memorize, self-improve, and find effective solutions. In this paper, we present a new learning paradigm, called Nested Learning (NL), that coherently represents a machine learning model with a set of nested, multi-level, and/or parallel optimization problems, each of which with its own context flow. Through the lenses of NL, existing deep learning methods learns from data through compressing their own context flow, and in-context learning naturally emerges in large models. NL suggests a philosophy to design more expressive learning algorithms with more levels, resulting in higher-order in-context learning and potentially unlocking effective continual learning capabilities. We advocate for NL by presenting three core contributions: (1) Expressive Optimizers: We show that known gradient-based optimizers, such as Adam, SGD with Momentum, etc., are in fact associative memory modules that aim to compress the gradients' information (by gradient descent). Building on this insight, we present other more expressive optimizers with deep memory and/or more powerful learning rules; (2) Self-Modifying Learning Module: Taking advantage of NL's insights on learning algorithms, we present a sequence model that learns how to modify itself by learning its own update algorithm; and (3) Continuum Memory System: We present a new formulation for memory system that generalizes the traditional viewpoint of long/short-term memory. Combining our self-modifying sequence model with the continuum memory system, we present a continual learning module, called Hope, showing promising results in language modeling, knowledge incorporation, and few-shot generalization tasks, continual learning, and long-context reasoning tasks.
T2TT3S9G
preprint
Sameera Ramasinghe, Thalaiyasingam Ajanthan, Gil Avraham, Yan Zuo, Alexander Long
2025
Saved 2026-02-19
Scaling models has led to significant advancements in deep learning, but training these models in decentralized settings remains challenging due to communication bottlenecks. While existing compression techniques are effective in data-parallel, they do not extend to model parallelism. Unlike data-parallel training, where weight gradients are exchanged, model-parallel requires compressing activations and activation gradients as they propagate through layers, accumulating compression errors. We propose a novel compression algorithm that compresses both forward and backward passes, enabling up to 99% compression with no convergence degradation with negligible memory/compute overhead. By leveraging a recursive structure in transformer networks, we predefine a low-dimensional subspace to confine the activations and gradients, allowing full reconstruction in subsequent layers. Our method achieves up to 100x improvement in communication efficiency and enables training billion-parameter-scale models over low-end GPUs connected via consumer-grade internet speeds as low as 80Mbps, matching the convergence of centralized datacenter systems with 100Gbps connections with model parallel.
7CUDNS6G
webpage
Sherry Yang, Yilun Du, Kamyar Ghasemipour, Jonathan Tompson, Leslie Kaelbling, Dale Schuurmans, Pieter Abbeel
2023
Saved 2026-02-18
Generative models trained on internet data have revolutionized how text, image, and video content can be created. Perhaps the next milestone for generative models is to simulate realistic experience in response to actions taken by humans, robots, and other interactive agents. Applications of a real-world simulator range from controllable content creation in games and movies, to training embodied agents purely in simulation that can be directly deployed in the real world. We explore the possibility of learning a universal simulator (UniSim) of real-world interaction through generative modeling. We first make the important observation that natural datasets available for learning a real-world simulator are often rich along different dimensions (e.g., abundant objects in image data, densely sampled actions in robotics data, and diverse movements in navigation data). With careful orchestration of diverse datasets, each providing a different aspect of the overall experience, we can simulate the visual outcome of both high-level instructions such as "open the drawer" and low-level controls from otherwise static scenes and objects. We use the simulator to train both high-level vision-language policies and low-level reinforcement learning policies, each of which can be deployed in the real world in zero shot after training purely in simulation. We also show that other types of intelligence such as video captioning models can benefit from training with simulated experience, opening up even wider applications. Video demos can be found at https://universal-simulator.github.io.
KC2FTHAI
journalArticle
A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27
Yann LeCun
Saved 2026-02-18
How could machines learn as efficiently as humans and animals? How could machines learn to reason and plan? How could machines learn representations of percepts and action plans at multiple levels of abstraction, enabling them to reason, predict, and plan at multiple time horizons? This position paper proposes an architecture and training paradigms with which to construct autonomous intelligent agents. It combines concepts such as configurable predictive world model, behavior driven through intrinsic motivation, and hierarchical joint embedding architectures trained with self-supervised learning.
HWZ3IP6F
preprint
Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, Gianluca Corrado
2023
Saved 2026-02-18
Autonomous driving promises transformative improvements to transportation, but building systems capable of safely navigating the unstructured complexity of real-world scenarios remains challenging. A critical problem lies in effectively predicting the various potential outcomes that may emerge in response to the vehicle's actions as the world evolves. To address this challenge, we introduce GAIA-1 ('Generative AI for Autonomy'), a generative world model that leverages video, text, and action inputs to generate realistic driving scenarios while offering fine-grained control over ego-vehicle behavior and scene features. Our approach casts world modeling as an unsupervised sequence modeling problem by mapping the inputs to discrete tokens, and predicting the next token in the sequence. Emerging properties from our model include learning high-level structures and scene dynamics, contextual awareness, generalization, and understanding of geometry. The power of GAIA-1's learned representation that captures expectations of future events, combined with its ability to generate realistic samples, provides new possibilities for innovation in the field of autonomy, enabling enhanced and accelerated training of autonomous driving technology.
V5HC4WQM
webpage
David Ha, Jürgen Schmidhuber
2018
Saved 2026-02-18
We explore building generative neural network models of popular reinforcement learning environments. Our world model can be trained quickly in an unsupervised manner to learn a compressed spatial and temporal representation of the environment. By using features extracted from the world model as inputs to an agent, we can train a very compact and simple policy that can solve the required task. We can even train our agent entirely inside of its own hallucinated dream generated by its world model, and transfer this policy back into the actual environment. An interactive version of this paper is available at https://worldmodels.github.io/
QBQ2LDMT
webpage
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap
2023
Saved 2026-02-18
Developing a general algorithm that learns to solve tasks across a wide range of applications has been a fundamental challenge in artificial intelligence. Although current reinforcement learning algorithms can be readily applied to tasks similar to what they have been developed for, configuring them for new application domains requires significant human expertise and experimentation. We present DreamerV3, a general algorithm that outperforms specialized methods across over 150 diverse tasks, with a single configuration. Dreamer learns a model of the environment and improves its behavior by imagining future scenarios. Robustness techniques based on normalization, balancing, and transformations enable stable learning across domains. Applied out of the box, Dreamer is the first algorithm to collect diamonds in Minecraft from scratch without human data or curricula. This achievement has been posed as a significant challenge in artificial intelligence that requires exploring farsighted strategies from pixels and sparse rewards in an open world. Our work allows solving challenging control problems without extensive experimentation, making reinforcement learning broadly applicable.
J982YRVF
blogPost
Jeffrey Wang
2026
Saved 2026-02-15
DCAI2P36
conferencePaper
Kenneth Li, Aspen K. Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, Martin Wattenberg
2022
Saved 2026-02-15
Language models show a surprising range of capabilities, but the source of their apparent competence is unclear. Do these networks just memorize a collection of surface statistics, or do they rely on internal representations of the process that generates the sequences they see? We investigate this question by applying a variant of the GPT model to the task of predicting legal moves in a simple board game, Othello. Although the network has no a priori knowledge of the game or its rules, we uncover evidence of an emergent nonlinear internal representation of the board state. Interventional experiments indicate this representation can be used to control the output of the network and create "latent saliency maps" that can help explain predictions in human terms.
UEHFSLT5
book
Algebra
Serge Lang
2005 · Springer
Saved 2026-02-13
JIGTTCU5
preprint
Joon Sung Park, Joseph C. O'Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, Michael S. Bernstein
2023
Saved 2026-02-12
Believable proxies of human behavior can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication to prototyping tools. In this paper, we introduce generative agents--computational software agents that simulate believable human behavior. Generative agents wake up, cook breakfast, and head to work; artists paint, while authors write; they form opinions, notice each other, and initiate conversations; they remember and reflect on days past as they plan the next day. To enable generative agents, we describe an architecture that extends a large language model to store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior. We instantiate generative agents to populate an interactive sandbox environment inspired by The Sims, where end users can interact with a small town of twenty five agents using natural language. In an evaluation, these generative agents produce believable individual and emergent social behaviors: for example, starting with only a single user-specified notion that one agent wants to throw a Valentine's Day party, the agents autonomously spread invitations to the party over the next two days, make new acquaintances, ask each other out on dates to the party, and coordinate to show up for the party together at the right time. We demonstrate through ablation that the components of our agent architecture--observation, planning, and reflection--each contribute critically to the believability of agent behavior. By fusing large language models with computational, interactive agents, this work introduces architectural and interaction patterns for enabling believable simulations of human behavior.
XHGPQRIW
preprint
Alex L. Zhang, Tim Kraska, Omar Khattab
2025
Saved 2026-02-12
We study allowing large language models (LLMs) to process arbitrarily long prompts through the lens of inference-time scaling. We propose Recursive Language Models (RLMs), a general inference strategy that treats long prompts as part of an external environment and allows the LLM to programmatically examine, decompose, and recursively call itself over snippets of the prompt. We find that RLMs successfully handle inputs up to two orders of magnitude beyond model context windows and, even for shorter prompts, dramatically outperform the quality of base LLMs and common long-context scaffolds across four diverse long-context tasks, while having comparable (or cheaper) cost per query.
YWE8TDZU
webpage
Saved 2026-02-12
QUYMMM5X
webpage
Saved 2026-02-12
Document preview for Milieudefensie et al. v. Royal Dutch Shell plc. - appeal
IGTHP6RI
webpage
Saved 2026-02-11
Genetically, I should be bald.    I started to lose my hair and go gray in my late 20s. Now, at 46, I’ve got a full head of hair and ~50% of my gray is gone. Here’s how I did it:   Start early and be proactive - I made the mistake of addressing my hair loss and graying after noticing it. By age 20, about 20% of men a
5TW2Q7L8
preprint
Jake Bruce, Michael Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar et al.
2024
Saved 2026-02-10
We introduce Genie, the first generative interactive environment trained in an unsupervised manner from unlabelled Internet videos. The model can be prompted to generate an endless variety of action-controllable virtual worlds described through text, synthetic images, photographs, and even sketches. At 11B parameters, Genie can be considered a foundation world model. It is comprised of a spatiotemporal video tokenizer, an autoregressive dynamics model, and a simple and scalable latent action model. Genie enables users to act in the generated environments on a frame-by-frame basis despite training without any ground-truth action labels or other domain-specific requirements typically found in the world model literature. Further the resulting learned latent action space facilitates training agents to imitate behaviors from unseen videos, opening the path for training generalist agents of the future.
DY2IEDPS
conferencePaper
Yajing Kong, Liu Liu, Jun Wang, Dacheng Tao
2021 · IEEE
Saved 2026-02-09
NP5LDKPX
preprint
Luke Darlow, Ciaran Regan, Sebastian Risi, Jeffrey Seely, Llion Jones
2025
Saved 2026-02-09
Biological brains demonstrate complex neural activity, where neural dynamics are critical to how brains process information. Most artificial neural networks ignore the complexity of individual neurons. We challenge that paradigm. By incorporating neuron-level processing and synchronization, we reintroduce neural timing as a foundational element. We present the Continuous Thought Machine (CTM), a model designed to leverage neural dynamics as its core representation. The CTM has two innovations: (1) neuron-level temporal processing, where each neuron uses unique weight parameters to process incoming histories; and (2) neural synchronization as a latent representation. The CTM aims to strike a balance between neuron abstractions and biological realism. It operates at a level of abstraction that effectively captures essential temporal dynamics while remaining computationally tractable. We demonstrate the CTM's performance and versatility across a range of tasks, including solving 2D mazes, ImageNet-1K classification, parity computation, and more. Beyond displaying rich internal representations and offering a natural avenue for interpretation owing to its internal process, the CTM is able to perform tasks that require complex sequential reasoning. The CTM can also leverage adaptive compute, where it can stop earlier for simpler tasks, or keep computing when faced with more challenging instances. The goal of this work is to share the CTM and its associated innovations, rather than pushing for new state-of-the-art results. To that end, we believe the CTM represents a significant step toward developing more biologically plausible and powerful artificial intelligence systems. We provide an accompanying interactive online demonstration at https://pub.sakana.ai/ctm/ and an extended technical report at https://pub.sakana.ai/ctm/paper .
QHXMYKPF
webpage
Jesper Ordrup
Saved 2026-02-09
A vocal technique reference covering 21 techniques: registers, styles, effects, embellish, and dynamics.
2DR69APH
webpage
Peter Steinberger
2025
Saved 2026-02-09
Why I stopped reading code and started watching it stream by.
8FVHLQPB
preprint
John Watrous
2025
Saved 2026-02-07
This is a course on the theory of quantum computing. It consists of 16 lessons, each with a video and written component, covering the basics of quantum information, quantum algorithms (including query algorithms, Shor's algorithm for integer factorization, and Grover's algorithm), the general formulation of quantum information (including density matrices, quantum channels, and general measurements), and quantum error correction (including the basics, the stabilizer formalism, CSS codes, the toric code, and fault-tolerant quantum computation).
XVFTHUZ4
webpage
The Founders' Tribune
2026
Saved 2026-02-07
Justine Musk is a Canadian author and ex-wife to Elon Musk.
XFUJ8PHS
blogPost
Saved 2026-02-06
Learn how to write killer sales emails at scale that engage buyers and drive meetings — including best practices for relevance, structure, concise messaging, Calls to Action (CTAs), and scaling outbound outreach without sacrificing quality.
W247ZKP2
journalArticle
Zhenyu Cai, Ryan Babbush, Simon C. Benjamin, Suguru Endo, William J. Huggins, Ying Li, Jarrod R. McClean, Thomas E. O’Brien
2023 · Reviews of Modern Physics
Saved 2026-02-06
J459YPPE
webpage
Saved 2026-02-05
What we learned from three iterations of a performance engineering take-home that Claude keeps beating.
XSTP5R84
preprint
Jingyu Liu, Xin Dong, Zhifan Ye, Rishabh Mehta, Yonggan Fu, Vartika Singh, Jan Kautz, Ce Zhang et al.
2025
Saved 2026-02-04
Diffusion language models hold the promise of fast parallel generation, while autoregressive (AR) models typically excel in quality due to their causal structure aligning naturally with language modeling. This raises a fundamental question: can we achieve a synergy with high throughput, higher GPU utilization, and AR level quality? Existing methods fail to effectively balance these two aspects, either prioritizing AR using a weaker model for sequential drafting (speculative decoding), leading to lower drafting efficiency, or using some form of left-to-right (AR-like) decoding logic for diffusion, which still suffers from quality degradation and forfeits its potential parallelizability. We introduce TiDAR, a sequence-level hybrid architecture that drafts tokens (Thinking) in Diffusion and samples final outputs (Talking) AutoRegressively - all within a single forward pass using specially designed structured attention masks. This design exploits the free GPU compute density, achieving a strong balance between drafting and verification capacity. Moreover, TiDAR is designed to be serving-friendly (low overhead) as a standalone model. We extensively evaluate TiDAR against AR models, speculative decoding, and diffusion variants across generative and likelihood tasks at 1.5B and 8B scales. Thanks to the parallel drafting and sampling as well as exact KV cache support, TiDAR outperforms speculative decoding in measured throughput and surpasses diffusion models like Dream and Llada in both efficiency and quality. Most notably, TiDAR is the first architecture to close the quality gap with AR models while delivering 4.71x to 5.91x more tokens per second.
X7UHC2JZ
videoRecording
Machine Learning Street Talk
2023
Saved 2026-02-03
Get 25% off Blinkist Annual premium! Start your 7-day free trial by clicking here: http://blinkist.com/machinelearningst... This is the first part of the Yannic Kilcher interview we posted recently! LAION petition mentioned: https://www.openpetition.eu/petition/... Yannic's channel: ‪@YannicKilcher‬ Tim Scarfe and Yannic Kilcher engaged in a wide-ranging discussion on the rise of powerful language models like ChatGPT and GPT-4 and their implications. While they believe these models will significantly transform industries and jobs, they do not foresee mass unemployment or existential risks from superintelligence. Scarfe and Kilcher think most people will adapt to work with AI, as they did with smartphones. The key will be learning to effectively apply these new tools. Regarding concerns about bias and misinformation, they acknowledge these are issues but not fundamentally new problems - people have always spread misinformation, and we must develop better digital literacy and detection methods. Regulating the technology itself likely will not solve underlying social challenges. Though skeptical of superintelligence, Scarfe and Kilcher expect models will become very capable in narrow domains, sometimes in ways beyond human understanding. But human judgment will remain critical to properly apply and constrain them. They believe progress should be encouraged if we are thoughtful and deliberate. Overall benefits likely far outweigh risks. Rather than limiting access, Scarfe and Kilcher support distributing models widely so researchers and users can better understand, improve and constrain them. Regulation and oversight still matter, but restricting progress seems misguided and unlikely to succeed. While powerful models may transform our lives, Scarfe and Kilcher expect human creativity and values to endure. With prudent development and policymaking, language models could usher in a new era of possibility. But we must be vigilant and deliberate to ensure the responsible and equitable progress of this technology. Overall, Scarfe and Kilcher express optimistic realism about the rise of language models. With open and proactively guided progress, they foresee emerging capabilities enhancing rather than diminishing human potential. But we must be willing partners in developing and applying these tools, not hapless bystanders. In shaping our future with AI, our values and judgment will be vital - these technologies remain a means, not an end. Their promise depends on human wisdom and oversight. TOC 0:00:00 - Introduction 0:03:09 - Background and Motivation 0:03:29 - AI Frameworks and Approaches 0:06:05 - Impact of Language Models 0:18:54 - Misinformation and Autonomy 0:29:43 - Superintelligence and AI Risk 0:45:15 - Discussion on the fear of superintelligence 0:51:37 - Understanding GPT-4 and its capabilities 0:55:55 - Comparing AI systems to human intelligence
K7D9U6EP
journalArticle
Welcome to a brave new world pitting processors against painters.
Joseph Pronechen
Saved 2026-02-03
QEMRWX4Z
webpage
Saved 2026-02-03
Manage your credits and payment history
AEC3MYAV
webpage
Saved 2026-02-01
DXQZGIKT
webpage
Saved 2026-01-30
9L7J27ND
journalArticle
Ryan LaRose, Andrea Mari, Sarah Kaiser, Peter J. Karalekas, Andre A. Alves, Piotr Czarnik, Mohamed El Mandouh, Max H. Gordon et al.
2022 · Quantum · Verein zur Förderung des Open Access Publizierens in den Quantenwissenschaften
Saved 2026-01-27
Ryan LaRose, Andrea Mari, Sarah Kaiser, Peter J. Karalekas, Andre A. Alves, Piotr Czarnik, Mohamed El Mandouh, Max H. Gordon, Yousef Hindy, Aaron Robertson, Purva Thakre, Misty Wahl, Danny Samuel, Rahul Mistri, Maxime Tremblay, Nick Gardner, Nathaniel T. Stemen, Nathan Shammah, and William J. Zeng, Quantum 6, 774 (2022). We introduce Mitiq, a Python package for error mitigation on noisy quantum computers. Error mitigation techniques can reduce the impact of noise on near-term quantum computers with minimal o…
595DP7KC
webpage
Saved 2026-01-27
8M22CFT9
document
Frobenius Normalization Enables Stable Training for Quantum State Denoising
Amitav Krishna
Saved 2026-01-27
SEEDZATB
journalArticle
Untitled
Saved 2026-01-27
BRQ9AD4Z
bookSection
Giacomo Mauro D’Ariano, Matteo G.A. Paris, Massimiliano F. Sacchi, Matteo Paris, Jaroslav Řeháček
2004 · Quantum State Estimation · Springer
Saved 2026-01-25
The state of a physical system is the mathematical object that provides a complete information on the system. The knowledge of the state is equivalent to know the result of any possible measurement on the system. This chapter reviews quantum state estimation for a generic quantum system by quantum tomography i.e. from the measurement of a suitable set of observables, a quorum, on repeated preparations of the system. Topics include characterization of quora, determination of the expectation value of any operator (including the nondiagonal projectors needed to construct a matrix representation of the density operator), evaluation of pattern functions, effect of instrumental noise, and example of tomographic procedure for harmonic systems and spins.
3CMC9VPY
bookSection
Matteo G.A. Paris, Jaroslav Řeháček, Matteo Paris, Jaroslav Řeháček
2004 · Quantum State Estimation · Springer
Saved 2026-01-25
The state of a physical system is the mathematical description of our knowledge of it, and provides information on its future and past. A state estimation technique is a method that provides the complete description of a system, i.e achieves the maximum possible knowledge of the state, thus allowing one to make the best, at least the best probabilistic, predictions on the results of any measurement that may be performed on the system.
TD9477UT
bookSection
Zdeněk Hradil, Jaroslav Řeháček, Jaromír Fiurášek, Miroslav Ježek, Matteo Paris, Jaroslav Řeháček
2004 · Quantum State Estimation · Springer
Saved 2026-01-25
Maximum Likelihood estimation is a versatile tool covering wide range of applications, but its benefits are apparent particularly in the quantum domain. For a given set of measurements, the most likely state is estimated. Though this problem is nonlinear, it can be effectively solved by an iterative algorithm exploiting the convexity of the likelihood functional and the manifold of density matrices. This formulation fully replaces the inverse Radon transformation routinely used for tomographic reconstructions. Moreover, it provides the most efficient estimation strategy saturating the Cramer-Rao lower bound asymptotically. In this sense it exploits the acquired data set in the optimal way and minimizes the artifacts associated with the reconstruction procedure. The idea of maximum likelihood reconstruction is further extended to the estimation of quantum processes, measurements, and discrimination between quantum states. This technique is well suited for future applications in quantum information science due to its ability to quantify very subtle and fragile quantum effects.
62V4C56Z
bookSection
Matteo Paris, Jaroslav Řeháček, Zdeněk Hradil, Jaroslav Řeháček, Jaromír Fiurášek, Miroslav Ježek
2004 · Quantum State Estimation · Springer Berlin Heidelberg
Saved 2026-01-25
Maximum Likelihood estimation is a versatile tool covering wide range of applications, but its benefits are apparent particularly in the quantum domain. For a given set of measurements, the most likely state is estimated. Though this problem is nonlinear, it can be effectively solved by an iterative algorithm exploiting the convexity of the likelihood functional and the manifold of density matrices. This formulation fully replaces the inverse Radon transformation routinely used for tomographic reconstructions. Moreover, it provides the most efficient estimation strategy saturating the Cramer-Rao lower bound asymptotically. In this sense it exploits the acquired data set in the optimal way and minimizes the artifacts associated with the reconstruction procedure. The idea of maximum likelihood reconstruction is further extended to the estimation of quantum processes, measurements, and discrimination between quantum states. This technique is well suited for future applications in quantum information science due to its ability to quantify very subtle and fragile quantum effects.
L8T38MEN
webpage
2026
Saved 2026-01-22
Dr. Crystal Senko Professor, Department of Physics and Astronomy Faculty of Science > Institute for Quantum Computing > Co-founder, Open Quantum Design Researchers from the University of Waterloo’s Faculty of Science and the Institute for Quantum Computing (IQC) are prioritizing collaboration over competition to advance quantum computer development and the field of quantum
7CTDM55U
conferencePaper
Xiao-Dao Lin, Hsi-Ming Chang, Jhih-Shih You, Hsiu-Chuan Hsu
2025 · 2025 IEEE International Conference on Quantum Control, Computing and Learning (qCCL)
Saved 2026-01-22
Quantum computing has gained significant attention in recent years, with numerous algorithms and applications under active development. Limited by the current quantum technology, quantum noise and readout error have become critical issues. Various methods have been proposed to address readout error through error mitigation techniques, typically involving post-processing of measurement data. However, most of these methods increase the quantum hardware overhead, leading to higher computational costs. In this work, we present a machine-learning-based approach that minimizes hardware overhead while improving accuracy of the measurement probability distributions. We employed a convolutional neural network (CNN) autoencoder, commonly used for image denoising, as our baseline model. The datasets were derived from 4-qubit random circuits with depths ranging from 1 to 18, generated using Qiskit backends for target and noisy measurement data. The model was trained using mean squared error (MSE) as the loss function and Adam optimizer over 500 epochs, achieving an average noise reduction by 95% across the validation set, with no signs of overfitting. To validate the model's effectiveness across diverse quantum states, we conducted extensive tests on both typical quantum circuits and algorithms, including Grover's search algorithm, Quantum Fourier Transform, Haar random circuits and Trivial Paramagnet. The results demonstrated consistent and robust denoising in noisy measurement data, indicating that the autoencoder model is well-suited for efficient quantum error mitigation for current noisy quantum computers. This work contributes to the advancement of quantum error mitigation techniques using machine learning.
AZ5GYTF5
journalArticle
Denoising weak lensing mass maps with diffusion model and generative adversarial network
Shohei D Aoyama, Ken Osato, Masato Shirasaki
Saved 2026-01-17
The matter distribution of the Universe can be mapped through the weak gravitational lensing (WL) effect: small distortions of the shapes of distant galaxies, which reflects the inhomogeneity of the cosmic density field. The most dominant contaminant in the WL effect is the shape noise; the signal is diluted due to the finite number of source galaxies. In order to explore the full potential of WL measurements, sharpening the signal by removing the shape noise from the observational data, i.e., WL denoising, is a pressing issue. Machine learning approaches, in particular, deep generative models, have proven effective at the WL denoising task. We implement a denoising model based on the diffusion model (DM) and conduct systematic in-depth comparisons with generative adversarial networks (GANs), which have been applied in previous works for WL denoising. Utilizing the large suite of mock simulations of WL observations, we demonstrate that DM surpasses GAN in the WL denosing task in multiple aspects: (1) the training process is more stable, (2) taking the average of multiple samples from DM can robustly reproduce the true signal, and (3) DM can recover various statistics with higher accuracy.
BAAR2Z7M
journalArticle
DiffICF: Diffusion-Driven Inverse Modeling for Laser Pulse Design in Inertial Confinement Fusion
Ricardo Luna Gutierrez, Vineet Gundecha, Rahman Ejaz, Varchas Gopalaswamy, Riccardo Betti, Sahand Ghorbanpour, Aarne Lees, Soumyendu Sarkar
Saved 2026-01-17
Traditional design of Laser Pulse Shapes (LPs) for Inertial Confinement Fusion (ICF) is a significant bottleneck, relying on computationally expensive simulations and manual iterative refinement. We introduce Diffusion-Driven Inverse Modeling for Laser Pulse Design (DiffICF), a generative inverse model that directly maps specified implosion outcomes to tailored LPs. DiffICF incorporates a physics-informed loss function that enforces known experimental and physical constraints. Moreover, it enables fine-grained control over pulse characteristics through constraint conditioning and inpainting. The efficacy of this framework was experimentally validated for optimizing implosion outcomes, offering a scalable, data-driven design tool to accelerate progress in fusion energy.
3BCPVYEN
preprint
Prithvi Raj
2026
Saved 2026-01-14
Learning an energy-based model (EBM) in the latent space of a top-down generative model offers a powerful framework for generation across many data modalities. However, it remains unclear how its interpretability can be used to guide model design, improve generative quality, and reduce training time. Moreover, the reliance on Langevin Monte Carlo (LMC) sampling presents challenges in efficiency and sampling multimodal latent distributions. We propose a novel adaptation of the Kolmogorov-Arnold representation theorem for generative modeling and introduce the Kolmogorov-Arnold Energy Model (KAEM) to take advantage of structural and inductive biases. By constraining the prior to univariate relationships, KAEM enables fast and exact inference via the inverse transform method. With the low dimensionality of the latent space and suitable inductive biases encoded, we demonstrate that importance sampling (IS) becomes a viable, unbiased, and highly efficient posterior sampler. For domains where IS fails, we introduce a strategy based on population-based LMC, decomposing the posterior into a sequence of annealed distributions to improve LMC mixing. KAEM balances common generative modeling trade-offs, offering fast inference, interpretability, and stable training, while being naturally suited to Zettascale Computing hardware.
QWK83Z8H
journalArticle
Sergey Bravyi, Andrew W. Cross, Jay M. Gambetta, Dmitri Maslov, Patrick Rall, Theodore J. Yoder
2024 · Nature · Nature Publishing Group
Saved 2026-01-11
The accumulation of physical errors1–3 prevents the execution of large-scale algorithms in current quantum computers. Quantum error correction4 promises a solution by encoding k logical qubits onto a larger number n of physical qubits, such that the physical errors are suppressed enough to allow running a desired computation with tolerable fidelity. Quantum error correction becomes practically realizable once the physical error rate is below a threshold value that depends on the choice of quantum code, syndrome measurement circuit and decoding algorithm5. We present an end-to-end quantum error correction protocol that implements fault-tolerant memory on the basis of a family of low-density parity-check codes6. Our approach achieves an error threshold of 0.7% for the standard circuit-based noise model, on par with the surface code7–10 that for 20 years was the leading code in terms of error threshold. The syndrome measurement cycle for a length-n code in our family requires n ancillary qubits and a depth-8 circuit with CNOT gates, qubit initializations and measurements. The required qubit connectivity is a degree-6 graph composed of two edge-disjoint planar subgraphs. In particular, we show that 12 logical qubits can be preserved for nearly 1 million syndrome cycles using 288 physical qubits in total, assuming the physical error rate of 0.1%, whereas the surface code would require nearly 3,000 physical qubits to achieve said performance. Our findings bring demonstrations of a low-overhead fault-tolerant quantum memory within the reach of near-term quantum processors.
HECRKJXM
journalArticle
The Pmarca Blog Archives
Marc Andreessen
Saved 2026-01-05
MJLN52FV
preprint
Christopher Woodyard
2026
Saved 2026-01-03
An exploratory, heuristic framework for interpreting the Riemann Hypothesis through the lens of recursive dynamics, signal processing, and geometric stability. Rather than proposing a formal proof, the work reframes fluctuations in the prime counting function as a normalized signal. It examines its behavior under a recursive projection operator inspired by resonance stabilization and fractal geometry. This paper also doesn't prove the Riemann Hypothesis, and I regret to inform the reader that no million-dollar check is attached.
D5JXGM92
journalArticle
Norhan Elsayed Amer, Walid Gomaa, Keiji Kimura, Kazunori Ueda, Ahmed El-Mahdy
2022 · EPJ Quantum Technology
Saved 2025-12-31
Abstract Current quantum processing technology is generally noisy with a limited number of qubits, stressing the importance of quantum state fidelity estimation. The complexity of this problem is mainly due to not only accounting for single gates and readout errors but also for interactions among which. Existing methods generally rely on either reconstructing the given circuit state, ideal state, and computing the distance of which; or forcing the system to be on a specific state. Both rely on conducting circuit measurements, in which computational efficiency is traded off with obtained fidelity details, requiring an exponential number of experiments for full information. This paper poses the question: Is the mapping between a given quantum circuit and its state fidelity learnable? If learnable, this would be a step towards an alternative approach that relies on machine learning, providing much more efficient computation. To answer this question, we propose three deep learning models for 1-, 3-, and 5-qubit circuits and experiment on the following real-quantum processors: ibmq_armonk (1-qubit), ibmq_lima (5-qubit) and ibmq_quito (5-qubit) backends, respectively. Our models achieved a mean correlation factor of 0.74, 0.67 and 0.66 for 1-, 3-, and 5-qubit random circuits, respectively, with the exponential state tomography method. Additionally, our 5-qubit model outperforms simple baseline state fidelity estimation method on three quantum benchmarks. Our method, trained on random circuits only, achieved a mean correlation factor of 0.968 while the baseline method achieved 0.738. Furthermore, we investigate the effect of dynamic noise on state fidelity estimation. The correlation factor substantially improved to 0.82 and 0.74 for the 3- and 5-qubit models, respectively. The results show that machine learning is promising for predicting state fidelity from circuit representation and this work may be considered a step towards efficient end-to-end learning.
JF6Z946Y
preprint
Changchun Feng, Laifa Tao, Lin Chen
2025
Saved 2025-12-31
Quantum state tomography (QST) faces exponential measurement requirements and noise sensitivity in multi-qubit systems, bottlenecking practical quantum technologies. We present a physics-informed neural network (PINN) framework integrating quantum mechanical constraints via adaptive weighting, a residual-and-attention-enhanced architecture, and differentiable Cholesky parameterization for physical validity. Evaluations on 2--5 qubit systems and arbitrary-dimensional states show PINN consistently outperforms traditional neural networks (TNNs), achieving highest fidelity across all dimensions. PINN outperforms baselines, with marked improvements in moderately high-dimensional systems, superior noise robustness (slower performance degradation), and consistent dimensional robustness. Theoretical analysis shows physical constraints reduce Rademacher complexity and mitigate the curse of dimensionality via constraint-induced dimension and sample complexity reduction, effective regardless of qubit number. While experiments are limited to 5-qubit systems due to computational constraints, our theoretical framework (convergence guarantees, generalization bounds, scalability theorems) justifies PINN's advantages will persist and strengthen in larger systems (6+ qubits), where constraint-induced dimension reduction benefits grow with system size. Practically, this advances quantum error correction and gate calibration by reducing measurement requirements from O(4^n) to O(2^n) while maintaining high fidelity, enabling faster error correction cycles and accelerated calibration critical for scalable quantum computing.
IYESEE6J
preprint
Daniel Uzcategui-Contreras, Antonio Guerra, Sebastian Niklitschek, Aldo Delgado
2025
Saved 2025-12-31
In this work, we propose a machine learning-based approach to address a specific aspect of the Quantum Marginal Problem: reconstructing a global density matrix compatible with a given set of quantum marginals. Our method integrates a quantum marginal imposition technique with convolutional denoising autoencoders. The loss function is carefully designed to enforce essential physical constraints, including Hermiticity, positivity, and normalization. Through extensive numerical simulations, we demonstrate the effectiveness of our approach, achieving high success rates and accuracy. Furthermore, we show that, in many cases, our model offers a faster alternative to state-of-the-art semidefinite programming solvers without compromising solution quality. These results highlight the potential of machine learning techniques for solving complex problems in quantum mechanics.
6P48WE9P
preprint
Yuchen Zhu, Tianrong Chen, Evangelos A. Theodorou, Xie Chen, Molei Tao
2024
Saved 2025-12-31
This article considers the generative modeling of the (mixed) states of quantum systems, and an approach based on denoising diffusion model is proposed. The key contribution is an algorithmic innovation that respects the physical nature of quantum states. More precisely, the commonly used density matrix representation of mixed-state has to be complex-valued Hermitian, positive semi-definite, and trace one. Generic diffusion models, or other generative methods, may not be able to generate data that strictly satisfy these structural constraints, even if all training data do. To develop a machine learning algorithm that has physics hard-wired in, we leverage mirror diffusion and borrow the physical notion of von Neumann entropy to design a new map, for enabling strict structure-preserving generation. Both unconditional generation and conditional generation via classifier-free guidance are experimentally demonstrated efficacious, the latter enabling the design of new quantum states when generated on unseen labels.
79Q92GXN
preprint
Mark Spivack, Orsola Rath Spivack
2024
Saved 2025-12-31
We discuss here the direct and inverse problems for wave propagation in a waveguide with rough internal surface and arbitrary mean shape. The high degree of multiple scattering inside the waveguide poses significant challenges both for the forward computation and for the recovery of the surface profile, and raises important questions about scattering cross-sections. This paper falls into two parts corresponding to these issues. We first apply Left-Right (L-R) operator splitting to calculate the scattered fields in 2- and 3-dimensional waveguides. Using this we illustrate the scattered fields and the effect of surface roughness. In the second part we formulate an algorithm for surface recovery from field measurements along the waveguide axis, which generalises recent work on surfaces in 2 dimensions. This method utilizes forward scattering assumptions in effect by formulating an integral equation in the unknown surface field, treated as a function of the surface. Although discussed in the context of waveguides, the formulae are given in a form applicable to a variety of geometries, with coefficients which will depend on and be determined by each specific application.
GL8TF8KK
book
Jake Skycak
Saved 2025-12-30
8F4QT9LX
book
Emperor of Rome Marcus Aurelius
2001
Saved 2025-12-30
KSKTML4N
preprint
Chaoyang He, Shen Li, Mahdi Soltanolkotabi, Salman Avestimehr
2021
Saved 2025-12-26
The size of Transformer models is growing at an unprecedented pace. It has only taken less than one year to reach trillion-level parameters after the release of GPT-3 (175B). Training such models requires both substantial engineering efforts and enormous computing resources, which are luxuries most research teams cannot afford. In this paper, we propose PipeTransformer, which leverages automated and elastic pipelining and data parallelism for efficient distributed training of Transformer models. PipeTransformer automatically adjusts the pipelining and data parallelism by identifying and freezing some layers during the training, and instead allocates resources for training of the remaining active layers. More specifically, PipeTransformer dynamically excludes converged layers from the pipeline, packs active layers into fewer GPUs, and forks more replicas to increase data-parallel width. We evaluate PipeTransformer using Vision Transformer (ViT) on ImageNet and BERT on GLUE and SQuAD datasets. Our results show that PipeTransformer attains a 2.4 fold speedup compared to the state-of-the-art baseline. We also provide various performance analyses for a more comprehensive understanding of our algorithmic and system-wise design. We also develop open-sourced flexible APIs for PipeTransformer, which offer a clean separation among the freeze algorithm, model definitions, and training accelerations, hence allowing it to be applied to other algorithms that require similar freezing strategies.
RC6RDGKP
preprint
Zhuohan Li, Siyuan Zhuang, Shiyuan Guo, Danyang Zhuo, Hao Zhang, Dawn Song, Ion Stoica
2021
Saved 2025-12-26
Model parallelism has become a necessity for training modern large-scale deep language models. In this work, we identify a new and orthogonal dimension from existing model parallel approaches: it is possible to perform pipeline parallelism within a single training sequence for Transformer-based language models thanks to its autoregressive property. This enables a more fine-grained pipeline compared with previous work. With this key idea, we design TeraPipe, a high-performance token-level pipeline parallel algorithm for synchronous model-parallel training of Transformer-based language models. We develop a novel dynamic programming-based algorithm to calculate the optimal pipelining execution scheme given a specific model and cluster configuration. We show that TeraPipe can speed up the training by 5.0x for the largest GPT-3 model with 175 billion parameters on an AWS cluster with 48 p3.16xlarge instances compared with state-of-the-art model-parallel methods. The code for reproduction can be found at https://github.com/zhuohan123/terapipe
PM4AH2SA
encyclopediaArticle
2025
Saved 2025-12-26
Raft is a consensus algorithm designed as an alternative to the Paxos family of algorithms. It was meant to be more understandable than Paxos by means of separation of logic, but it is also formally proven safe and offers some additional features. Raft offers a generic way to distribute a state machine across a cluster of computing systems, ensuring that each node in the cluster agrees upon the same series of state transitions. It has a number of open-source reference implementations, with full-specification implementations in Go, C++, Java, JavaScript, and Scala. It is named after Reliable, Replicated, Redundant, And Fault-Tolerant. Raft is not Byzantine fault tolerant; the nodes trust the elected leader, and the algorithm assumes all participants are trustworthy.
FMEHWMY7
encyclopediaArticle
2025
Saved 2025-12-26
In computer science, Paxos is a family of protocols for solving consensus in a network of unreliable or fallible processors. Consensus is the process of agreeing on one result among a group of participants. This problem becomes difficult when the participants or their communications may experience failures. Consensus protocols are the basis for the state machine replication approach to distributed computing, as suggested by Leslie Lamport and surveyed by Fred Schneider. State machine replication is a technique for converting an algorithm into a fault-tolerant, distributed implementation. Ad-hoc techniques may leave important cases of failures unresolved. The principled approach proposed by Lamport et al. ensures all cases are handled safely. The Paxos protocol was first submitted in 1989 and named after a fictional legislative consensus system used on the Paxos island in Greece, where Lamport wrote that the parliament had to function "even though legislators continually wandered in and out of the parliamentary Chamber". It was later published as a journal article in 1998. The Paxos family of protocols includes a spectrum of trade-offs between the number of processors, number of message delays before learning the agreed value, the activity level of individual participants, number of messages sent, and types of failures. Although no deterministic fault-tolerant consensus protocol can guarantee progress in an asynchronous network (a result proved in a paper by Fischer, Lynch and Paterson), Paxos guarantees safety (consistency), and the conditions that could prevent it from making progress are difficult to provoke. Paxos is usually used where durability is required (for example, to replicate a file or a database), in which the amount of durable state could be large. The protocol attempts to make progress even during periods when some bounded number of replicas are unresponsive. There is also a mechanism to drop a permanently failed replica or to add a new replica.
S8IH3YX7
encyclopediaArticle
2025
Saved 2025-12-26
In cryptography, a zero-knowledge proof (also known as a ZK proof or ZKP) is a protocol in which one party (the prover) can convince another party (the verifier) that some given statement is true, without conveying to the verifier any information beyond the mere fact of that statement's truth. The intuition behind the nontriviality of zero-knowledge proofs is that it is trivial to prove possession of the relevant information simply by revealing it; the hard part is to prove this possession without revealing this information (or any aspect of it whatsoever). In light of the fact that one should be able to generate a proof of some statement only when in possession of certain secret information connected to the statement, the verifier, even after having become convinced of the statement's truth by means of a zero-knowledge proof, should nonetheless remain unable to prove the statement to further third parties. Zero-knowledge proofs can be interactive, meaning that the prover and verifier exchange messages according to some protocol, or noninteractive, meaning that the verifier is convinced by a single prover message and no other communication is needed. In the standard model, interaction is required, except for trivial proofs of BPP problems. In the common random string and random oracle models, non-interactive zero-knowledge proofs exist. The Fiat–Shamir heuristic can be used to transform certain interactive zero-knowledge proofs into noninteractive ones.
IJKSZ2U6
webpage
Saved 2025-12-26
YB6JB4CW
encyclopediaArticle
2025
Saved 2025-12-26
A cryptographic hash function (CHF) is a hash algorithm (a map of an arbitrary binary string to a binary string with a fixed size of n {\displaystyle n} bits) that has special properties desirable for a cryptographic application: the probability of a particular n {\displaystyle n} -bit output result (hash value) for a random input string ("message") is 2 − n {\displaystyle 2^{-n}} (as for any good hash), so the hash value can be used as a representative of the message; finding an input string that matches a given hash value (a pre-image) is infeasible, assuming all input strings are equally likely. The resistance to such search is quantified as security strength: a cryptographic hash with n {\displaystyle n} bits of hash value is expected to have a preimage resistance strength of n {\displaystyle n} bits, unless the space of possible input values is significantly smaller than 2 n {\displaystyle 2^{n}} (a practical example can be found in § Attacks on hashed passwords); a second preimage resistance strength, with the same expectations, refers to a similar problem of finding a second message that matches the given hash value when one message is already known; finding any pair of different messages that yield the same hash value (a collision) is also infeasible: a cryptographic hash is expected to have a collision resistance strength of n / 2 {\displaystyle n/2} bits (lower because of the birthday paradox). Cryptographic hash functions have many information-security applications, notably in digital signatures, message authentication codes (MACs), and other forms of authentication. They can also be used as ordinary hash functions, to index data in hash tables, for fingerprinting, to detect duplicate data or uniquely identify files, and as checksums to detect accidental data corruption. Indeed, in information-security contexts, cryptographic hash values are sometimes called (digital) fingerprints, checksums, (message) digests, or just hash values, even though all these terms stand for more general functions with rather different properties and purposes. Non-cryptographic hash functions are used in hash tables and to detect accidental errors; their constructions frequently provide no resistance to a deliberate attack. For example, a denial-of-service attack on hash tables is possible if the collisions are easy to find, as in the case of linear cyclic redundancy check (CRC) functions.
UCI93ZR9
encyclopediaArticle
2025
Saved 2025-12-26
In cryptography and computer science, a hash tree or Merkle tree is a tree in which every "leaf" node is labelled with the cryptographic hash of a data block, and every node that is not a leaf (called a branch, inner node, or inode) is labelled with the cryptographic hash of the labels of its child nodes. A hash tree allows efficient and secure verification of the contents of a large data structure. A hash tree is a generalization of a hash list and a hash chain. Demonstrating that a leaf node is a part of a given binary hash tree requires computing a number of hashes proportional to the logarithm of the number of leaf nodes in the tree. Conversely, in a hash list, the number is proportional to the number of leaf nodes itself. A Merkle tree is therefore an efficient example of a cryptographic commitment scheme, in which the root of the tree is seen as a commitment and leaf nodes may be revealed and proven to be part of the original commitment. The concept of a hash tree is named after Ralph Merkle, who patented it in 1979.
GTJHQF46
encyclopediaArticle
2025
Saved 2025-12-26
A Byzantine fault is a condition of a system, particularly a distributed computing system, where a fault occurs such that different symptoms are presented to different observers, including imperfect information on whether a system component has failed. The term takes its name from an allegory, the "Byzantine generals problem", developed to describe a situation in which, to avoid catastrophic failure of a system, the system's actors must agree on a strategy, but some of these actors are unreliable in such a way as to cause other (good) actors to disagree on the strategy and they may be unaware of the disagreement. A Byzantine fault is also known as a Byzantine generals problem, a Byzantine agreement problem, or a Byzantine failure. Byzantine fault tolerance (BFT) is the resilience of a fault-tolerant computer system or similar system to such conditions.
LAM4ZLZH
computerProgram
2025
Saved 2025-12-26
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
V4FANARD
webpage
Michael Goin
2025
Saved 2025-12-26
Explore how distributed inference works within vLLM in this recap of Neural Magic's vLLM Office Hours with Michael Goin and Murali Andoorveedu, a vLLM committer from CentML
SCKUIQNP
blogPost
Sarat Kannan
2025
Saved 2025-12-26
Introduction
UZL3NY3J
webpage
Saved 2025-12-26
Large language models are among the most significant recent advances in machine learning. Still, leveraging these models can be difficult: offloading and quantization have limitations, and third-party APIs are less flexible. We propose Petals, an open-source decentralized system (showcased this week at the ACL 2023 Demonstrations track) allowing anybody to run large models or even adapt them using the idle resources of volunteers. In this post, you will learn the motivation behind the system, its underlying ideas, and its advantages compared to other ways of using large models.
VBX34V9Y
preprint
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, Bryan Catanzaro
2020
Saved 2025-12-26
Recent work in language modeling demonstrates that training large transformer models advances the state of the art in Natural Language Processing applications. However, very large models can be quite difficult to train due to memory constraints. In this work, we present our techniques for training very large transformer models and implement a simple, efficient intra-layer model parallel approach that enables training transformer models with billions of parameters. Our approach does not require a new compiler or library changes, is orthogonal and complimentary to pipeline model parallelism, and can be fully implemented with the insertion of a few communication operations in native PyTorch. We illustrate this approach by converging transformer based models up to 8.3 billion parameters using 512 GPUs. We sustain 15.1 PetaFLOPs across the entire application with 76% scaling efficiency when compared to a strong single GPU baseline that sustains 39 TeraFLOPs, which is 30% of peak FLOPs. To demonstrate that large language models can further advance the state of the art (SOTA), we train an 8.3 billion parameter transformer language model similar to GPT-2 and a 3.9 billion parameter model similar to BERT. We show that careful attention to the placement of layer normalization in BERT-like models is critical to achieving increased performance as the model size grows. Using the GPT-2 model we achieve SOTA results on the WikiText103 (10.8 compared to SOTA perplexity of 15.8) and LAMBADA (66.5% compared to SOTA accuracy of 63.2%) datasets. Our BERT model achieves SOTA results on the RACE dataset (90.9% compared to SOTA accuracy of 89.4%).
PCQWDVHK
videoRecording
Andrej Karpathy
2023
Saved 2025-12-26
We build a Generatively Pretrained Transformer (GPT), following the paper "Attention is All You Need" and OpenAI's GPT-2 / GPT-3. We talk about connections to ChatGPT, which has taken the world by storm. We watch GitHub Copilot, itself a GPT, help us write a GPT (meta :D!) . I recommend people watch the earlier makemore videos to get comfortable with the autoregressive language modeling framework and basics of tensors and PyTorch nn, which we take for granted in this video.
7X9A3TGG
webpage
Jay Alammar
Saved 2025-12-26
Discussions: Hacker News (65 points, 4 comments), Reddit r/MachineLearning (29 points, 3 comments) Translations: Arabic, Chinese (Simplified) 1, Chinese (Simplified) 2, French 1, French 2, Italian, Japanese, Korean, Persian, Russian, Spanish 1, Spanish 2, Vietnamese Watch: MIT’s Deep Learning State of the Art lecture referencing this post Featured in courses at Stanford, Harvard, MIT, Princeton, CMU and others Update: This post has now become a book! Check out LLM-book.com which contains (Chapter 3) an updated and expanded version of this post speaking about the latest Transformer models and how they've evolved in the seven years since the original Transformer (like Multi-Query Attention and RoPE Positional embeddings). In the previous post, we looked at Attention – a ubiquitous method in modern deep learning models. Attention is a concept that helped improve the performance of neural machine translation applications. In this post, we will look at The Transformer – a model that uses attention to boost the speed with which these models can be trained. The Transformer outperforms the Google Neural Machine Translation model in specific tasks. The biggest benefit, however, comes from how The Transformer lends itself to parallelization. It is in fact Google Cloud’s recommendation to use The Transformer as a reference model to use their Cloud TPU offering. So let’s try to break the model apart and look at how it functions. The Transformer was proposed in the paper Attention is All You Need. A TensorFlow implementation of it is available as a part of the Tensor2Tensor package. Harvard’s NLP group created a guide annotating the paper with PyTorch implementation. In this post, we will attempt to oversimplify things a bit and introduce the concepts one by one to hopefully make it easier to understand to people without in-depth knowledge of the subject matter. 2025 Update: We’ve built a free short course that brings the contents of this post up-to-date with animations: A High-Level Look Let’s begin by looking at the model as a single black box. In a machine translation application, it would take a sentence in one language, and output its translation in another.
FEQY4WBC
journalArticle
Figure 1: the cover of Knuth’s letter
P van Emde Boas
Saved 2025-12-23
FRDD4Z5X
book
The gay science: with a prelude in German rhymes and an appendix of songs
Friedrich Nietzsche, Bernard Williams, Friedrich Nietzsche
2004 · Cambridge University Press
Saved 2025-12-23
K453VWJL
journalArticle
Beyond Good and Evil: Prelude to a Philosophy of the Future
Friedrich Nietzsche
· HISTORY OF PHILOSOPHY
Saved 2025-12-23
2YTTWV7U
preprint
Zhangde Song, Jieyu Lu, Yuanqi Du, Botao Yu, Thomas M. Pruyn, Yue Huang, Kehan Guo, Xiuzhe Luo et al.
2025
Saved 2025-12-23
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific "superintelligence". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.
Z5TMF9HY
preprint
Ali Behrouz, Peilin Zhong, Vahab Mirrokni
2024
Saved 2025-12-22
Over more than a decade there has been an extensive research effort on how to effectively utilize recurrent models and attention. While recurrent models aim to compress the data into a fixed-size memory (called hidden state), attention allows attending to the entire context window, capturing the direct dependencies of all tokens. This more accurate modeling of dependencies, however, comes with a quadratic cost, limiting the model to a fixed-length context. We present a new neural long-term memory module that learns to memorize historical context and helps attention to attend to the current context while utilizing long past information. We show that this neural memory has the advantage of fast parallelizable training while maintaining a fast inference. From a memory perspective, we argue that attention due to its limited context but accurate dependency modeling performs as a short-term memory, while neural memory due to its ability to memorize the data, acts as a long-term, more persistent, memory. Based on these two modules, we introduce a new family of architectures, called Titans, and present three variants to address how one can effectively incorporate memory into this architecture. Our experimental results on language modeling, common-sense reasoning, genomics, and time series tasks show that Titans are more effective than Transformers and recent modern linear recurrent models. They further can effectively scale to larger than 2M context window size with higher accuracy in needle-in-haystack tasks compared to baselines.
9EUPIKEQ
preprint
Igor Shilov, Alex Cloud, Aryo Pradipta Gema, Jacob Goldman-Wetzler, Nina Panickssery, Henry Sleight, Erik Jones, Cem Anil
2025
Saved 2025-12-20
Large Language Models increasingly possess capabilities that carry dual-use risks. While data filtering has emerged as a pretraining-time mitigation, it faces significant challenges: labeling whether data is harmful is expensive at scale, and given improving sample efficiency with larger models, even small amounts of mislabeled content could give rise to dangerous capabilities. To address risks associated with mislabeled harmful content, prior work proposed Gradient Routing (Cloud et al., 2024) -- a technique that localizes target knowledge into a dedicated subset of model parameters so they can later be removed. We explore an improved variant of Gradient Routing, which we call Selective GradienT Masking (SGTM), with particular focus on evaluating its robustness to label noise. SGTM zero-masks selected gradients such that target domain examples only update their dedicated parameters. We test SGTM's effectiveness in two applications: removing knowledge of one language from a model trained on a bilingual synthetic dataset, and removing biology knowledge from a model trained on English Wikipedia. In both cases SGTM provides better retain/forget trade-off in the presence of labeling errors compared to both data filtering and a previously proposed instantiation of Gradient Routing. Unlike shallow unlearning approaches that can be quickly undone through fine-tuning, SGTM exhibits strong robustness to adversarial fine-tuning, requiring seven times more fine-tuning steps to reach baseline performance on the forget set compared to a finetuning-based unlearning method (RMU). Our results suggest SGTM provides a promising pretraining-time complement to existing safety mitigations, particularly in settings where label noise is unavoidable.
KUSWJCN4
journalArticle
Christopher Chamberland, Pooya Ronagh
2018 · Quantum Science and Technology
Saved 2025-12-19
C2NEFSWY
journalArticle
Nikolas P. Breuckmann, Xiaotong Ni
2018 · Quantum
Saved 2025-12-19
Machine learning has the potential to become an important tool in quantum error correction as it allows the decoder to adapt to the error distribution of a quantum chip. An additional motivation for using neural networks is the fact that they can be evaluated by dedicated hardware which is very fast and consumes little power. Machine learning has been previously applied to decode the surface code. However, these approaches are not scalable as the training has to be redone for every system size which becomes increasingly difficult. In this work the existence of local decoders for higher dimensional codes leads us to use a low-depth convolutional neural network to locally assign a likelihood of error on each qubit. For noiseless syndrome measurements, numerical simulations show that the decoder has a threshold of around 7.1% when applied to the 4D toric code. When the syndrome measurements are noisy, the decoder performs better for larger code sizes when the error probability is low. We also give theoretical and numerical analysis to show how a convolutional neural network is different from the 1-nearest neighbor algorithm, which is a baseline machine learning method.
J6HHSW78
preprint
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, Illia Polosukhin
2023
Saved 2025-12-18
The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely. Experiments on two machine translation tasks show these models to be superior in quality while being more parallelizable and requiring significantly less time to train. Our model achieves 28.4 BLEU on the WMT 2014 English-to-German translation task, improving over the existing best results, including ensembles by over 2 BLEU. On the WMT 2014 English-to-French translation task, our model establishes a new single-model state-of-the-art BLEU score of 41.8 after training for 3.5 days on eight GPUs, a small fraction of the training costs of the best models from the literature. We show that the Transformer generalizes well to other tasks by applying it successfully to English constituency parsing both with large and limited training data.
RIA9ZT3P
webpage
Saved 2025-12-15
E953VIFC
blogPost
2024
Saved 2025-12-15
While PC and console may be growing, playtime is declining. This article gives you a glimpse into how our new report covers gaming engagement.
WI85EJ4Q
webpage
2023
Saved 2025-12-15
Interesting Steam user stats. Includes data on users, games, and more.
738FZI7Q
blogPost
2025
Saved 2025-12-15
This article highlights the key global games market estimates and forecasts for 2025 and what’s ahead in the years to come.
FBI4I222
preprint
Alexander Borzunov, Dmitry Baranchuk, Tim Dettmers, Max Ryabinin, Younes Belkada, Artem Chumachenko, Pavel Samygin, Colin Raffel
2023
Saved 2025-12-15
Many NLP tasks benefit from using large language models (LLMs) that often have more than 100 billion parameters. With the release of BLOOM-176B and OPT-175B, everyone can download pretrained models of this scale. Still, using these models requires highend hardware unavailable to many researchers. In some cases, LLMs can be used more affordably via RAM offloading or hosted APIs. However, these techniques have innate limitations: offloading is too slow for interactive inference, while APIs are not flexible enough for research that requires access to weights, attention or logits. In this work, we propose PETALS1 — a system for inference and fine-tuning of large models collaboratively by joining the resources of multiple parties. We demonstrate that this strategy outperforms offloading for very large models, running inference of BLOOM-176B on consumer GPUs with ≈ 1 step per second, which is enough for many interactive LLM applications. Unlike most inference APIs, PETALS also natively exposes hidden states of served models, allowing to train and share custom model extensions based on efficient fine-tuning methods.
GB7HEE7Q
preprint
Alexander Borzunov, Dmitry Baranchuk, Tim Dettmers, Max Ryabinin, Younes Belkada, Artem Chumachenko, Pavel Samygin, Colin Raffel
2023
Saved 2025-12-15
Many NLP tasks benefit from using large language models (LLMs) that often have more than 100 billion parameters. With the release of BLOOM-176B and OPT-175B, everyone can download pretrained models of this scale. Still, using these models requires highend hardware unavailable to many researchers. In some cases, LLMs can be used more affordably via RAM offloading or hosted APIs. However, these techniques have innate limitations: offloading is too slow for interactive inference, while APIs are not flexible enough for research that requires access to weights, attention or logits. In this work, we propose PETALS1 — a system for inference and fine-tuning of large models collaboratively by joining the resources of multiple parties. We demonstrate that this strategy outperforms offloading for very large models, running inference of BLOOM-176B on consumer GPUs with ≈ 1 step per second, which is enough for many interactive LLM applications. Unlike most inference APIs, PETALS also natively exposes hidden states of served models, allowing to train and share custom model extensions based on efficient fine-tuning methods.
D3J4VISB
journalArticle
Giacomo Torlai, Guglielmo Mazzola, Juan Carrasquilla, Matthias Troyer, Roger Melko, Giuseppe Carleo
2018 · Nature Physics
Saved 2025-12-12
The experimental realization of increasingly complex synthetic quantum systems calls for the development of general theoretical methods to validate and fully exploit quantum resources. Quantum state tomography (QST) aims to reconstruct the full quantum state from simple measurements, and therefore provides a key tool to obtain reliable analytics1–3. However, exact brute-force approaches to QST place a high demand on computational resources, making them unfeasible for anything except small systems4,5. Here we show how machine learning techniques can be used to perform QST of highly entangled states with more than a hundred qubits, to a high degree of accuracy. We demonstrate that machine learning allows one to reconstruct traditionally challenging many-body quantities—such as the entanglement entropy—from simple, experimentally accessible measurements. This approach can benefit existing and future generations of devices ranging from quantum computers to ultracold-atom quantum simulators6–8.
9F6E2AXX
journalArticle
Dominik Koutný, Libor Motka, Zdeněk Hradil, Jaroslav Řeháček, Luis L. Sánchez-Soto
2022 · Physical Review A
Saved 2025-12-12
2KVNZ8AP
preprint
Preetum Nakkiran, Arwen Bradley, Hattie Zhou, Madhu Advani
2024
Saved 2025-12-11
We present an accessible first course on diffusion models and flow matching for machine learning, aimed at a technical audience with no diffusion experience. We try to simplify the mathematical details as much as possible (sometimes heuristically), while retaining enough precision to derive correct algorithms.
VE6U4S2N
journalArticle
Angela Rosy Morgillo, Stefano Mangini, Marco Piastra, Chiara Macchiavello
2024 · Quantum Machine Intelligence
Saved 2025-12-11
Quantum noise is currently limiting efficient quantum information processing and computation, impacting on the fidelity and reliability of quantum states. In this work, we consider the tasks of reconstructing and classifying quantum states corrupted by the action of an unknown noisy channel using classical feed-forward neural networks. By framing reconstruction as a regression problem, we show how such an approach can be used to recover with fidelities exceeding 99% the noiseless density matrices of quantum states of up to three qubits undergoing noisy evolution, and we test its performance with both single-qubit (bit-flip, phase-flip, depolarizing, and amplitude damping) and two-qubit quantum channels (correlated amplitude damping). Furthermore, a critical aspect of our investigation involves also a comprehensive comparison between mean squared error and infidelity as loss functions. Our findings reveal that these two metrics yield comparable results in the context of state reconstruction. Moreover, we also consider the task of distinguishing between different quantum noisy channels, and show how a neural network-based classifier is able to solve such a classification problem with perfect accuracy.
CG7YKXNI
book
Quantum Computation and Quantum Information
Micheal Nielsen
· 0-521-63503-9
Saved 2025-12-09
GRDYYE9Y
webpage
Paul Graham
Saved 2025-12-09
79DJP64N
blogPost
Saved 2025-12-08
Quantum is an open-access peer-reviewed journal for quantum science and related fields. Quantum is non-profit and community-run: an effort by researchers and for researchers to make science more open and publishing more transparent and efficient.
FA9VPU95
journalArticle
Daniel Bultrini, Max Hunter Gordon, Piotr Czarnik, Andrew Arrasmith, M. Cerezo, Patrick J. Coles, Lukasz Cincio
2023 · Quantum · Verein zur Förderung des Open Access Publizierens in den Quantenwissenschaften
Saved 2025-12-08
Daniel Bultrini, Max Hunter Gordon, Piotr Czarnik, Andrew Arrasmith, M. Cerezo, Patrick J. Coles, and Lukasz Cincio, Quantum 7, 1034 (2023). Error mitigation is an essential component of achieving a practical quantum advantage in the near term, and a number of different approaches have been proposed. In this work, we recognize th…
SP9JHZ6N
preprint
Xiao-Yue Xu, Xin Xue, Tianyu Chen, Chen Ding, Tian Li, Haoyi Zhou, He-Liang Huang, Wan-Su Bao
2025
Saved 2025-12-08
Noise is a major obstacle in current quantum computing, and Machine Learning for Quantum Error Mitigation (ML-QEM) promises to address this challenge, enhancing computational accuracy while reducing the sampling overheads of standard QEM methods. Yet, existing models lack physical interpretability and rely heavily on extensive datasets, hindering their scalability in large-scale quantum circuits. To tackle these issues, we introduce the Neural Noise Accumulation Surrogate (NNAS), a physics-inspired neural network for ML-QEM that incorporates the structural characteristics of quantum noise accumulation within multi-layer circuits, endowing the model with physical interpretability. Experimental results demonstrate that NNAS outperforms current methods across a spectrum of metrics, including error mitigation capability, quantum resource consumption, and training dataset size. Notably, for deeper circuits where QEM methods typically struggle, NNAS achieves a remarkable reduction of over half in errors. NNAS also demands substantially fewer training data, reducing dataset reliance by at least an order of magnitude, due to its ability to rapidly capture noise accumulation patterns across circuit layers. This work pioneers the integration of quantum process-derived structural characteristics into neural network architectures, broadly enhancing QEM's performance and applicability, and establishes an integrative paradigm that extends to various quantum-inspired neural network architectures.
N25EJVHX
journalArticle
Lukasz Cincio, Kenneth Rudinger, Mohan Sarovar, Patrick J. Coles
2021 · PRX Quantum · American Physical Society
Saved 2025-12-08
Noise mitigation and reduction will be crucial for obtaining useful answers from near-term quantum computers. In this work, we present a general framework based on machine learning for reducing the impact of quantum hardware noise on quantum circuits. Our method, called noise-aware circuit learning (NACL), applies to circuits designed to compute a unitary transformation, prepare a set of quantum states, or estimate an observable of a many-qubit state. Given a task and a device model that captures information about the noise and connectivity of qubits in a device, NACL outputs an optimized circuit to accomplish this task in the presence of noise. It does so by minimizing a task-specific cost function over circuit depths and circuit structures. To demonstrate NACL, we construct circuits resilient to a fine-grained noise model derived from gate set tomography on a superconducting-circuit quantum device, for applications including quantum state overlap, quantum Fourier transform, and 𝑊-state preparation.
CRPJTWUN
journalArticle
Rajeev Acharya, Dmitry A. Abanin, Laleh Aghababaie-Beni, Igor Aleiner, Trond I. Andersen, Markus Ansmann, Frank Arute, Kunal Arya et al.
2024 · Nature · Springer Science and Business Media LLC
Saved 2025-12-07
72W9HRV8
journalArticle
Sergey Bravyi, Andrew W. Cross, Jay M. Gambetta, Dmitri Maslov, Patrick Rall, Theodore J. Yoder
2024 · Nature · Springer Science and Business Media LLC
Saved 2025-12-07
SU9SSFJX
document
Sergey N. Filippov, Sabrina Maniscalco, Guillermo García-Pérez
2024
Saved 2025-12-07
3H9SKLFG
journalArticle
Zhenyu Cai, Ryan Babbush, Simon C. Benjamin, Suguru Endo, William J. Huggins, Ying Li, Jarrod R. McClean, Thomas E. O'Brien
2023 · Rev. Mod. Phys. · American Physical Society
Saved 2025-12-07
ZPMVBYG7
journalArticle
David F. Locher, Lorenzo Cardarelli, Markus Müller
2023 · Quantum · Verein zur Förderung des Open Access Publizierens in den Quantenwissenschaften
Saved 2025-12-07
MPSSKCNI
journalArticle
Angela Rosy Morgillo, Stefano Mangini, Marco Piastra, Chiara Macchiavello
2024 · Quantum Machine Intelligence · Springer Science and Business Media LLC
Saved 2025-12-07
2UMS87NR
journalArticle
Zonghan Wu, Shirui Pan, Fengwen Chen, Guodong Long, Chengqi Zhang, Philip S. Yu
2021 · IEEE Transactions on Neural Networks and Learning Systems · Institute of Electrical and Electronics Engineers (IEEE)
Saved 2025-12-07
NQDIGPRS
journalArticle
Xuanqiang Zhao, Benchi Zhao, Zihan Xia, Xin Wang
2023 · Quantum · Verein zur Forderung des Open Access Publizierens in den Quantenwissenschaften
Saved 2025-12-07
XLXR48Z5
journalArticle
Giulia Marcucci, Davide Pierangeli, Claudio Conti
2020 · Phys. Rev. Lett. · American Physical Society
Saved 2025-12-07
FVI9C3KY
document
Hayden Joy, Marios Mattheakis, Pavlos Protopapas
2022
Saved 2025-12-07
TUFL2DVH
document
Marios Mattheakis, Gabriel R. Schleder, Daniel T. Larson, Efthimios Kaxiras
2022
Saved 2025-12-07
QN7AYEUB
document
Shaan Desai, Marios Mattheakis, Hayden Joy, Pavlos Protopapas, Stephen Roberts
2022
Saved 2025-12-07
FQGMYPGT
document
M. Mattheakis, P. Protopapas, D. Sondak, M. Di Giovanni, E. Kaxiras
2020
Saved 2025-12-07
CEJA7ED4
journalArticle
Marios Mattheakis, David Sondak, Akshunna S. Dogra, Pavlos Protopapas
2022 · Physical Review E · American Physical Society (APS)
Saved 2025-12-07
HNELZ2KZ
journalArticle
Feiyu Chen, David Sondak, Pavlos Protopapas, Marios Mattheakis, Shuheng Liu, Devansh Agarwal, Marco Di Giovanni
2020 · Journal of Open Source Software · The Open Journal
Saved 2025-12-07
26LZ9C7A
book
Practical Common Lisp
Peter Seibel
2005
Saved 2025-12-07
TGKBC27F
book
Common Lisp Cookbook
Vincent Dardel
2022
Saved 2025-12-07
2M63PUJG
journalArticle
Cate Hall
2023 · Useful Fictions
Saved 2025-12-07
I5N3ZZ57
journalArticle
Bo Qi, Zhibo Hou, Li Li, Daoyi Dong, Guoyong Xiang, Guangcan Guo
2013 · Scientific Reports · Springer Science and Business Media LLC
Saved 2025-12-07
N3R68J52
journalArticle
When quantum state tomography benefits from willful ignorance
Libor Motka, Martin Paúr, Jaroslav Rehacek, Zdenek Hradil, L. Sanchez-Soto
2021 · New Journal of Physics
Saved 2025-12-07
7H2LN5YV
conferencePaper
Ryan O'Donnell, John Wright
2016 · Proceedings of the Forty-Eighth Annual ACM Symposium on Theory of Computing · Association for Computing Machinery
Saved 2025-12-07
In the quantum state tomography problem, one wishes to estimate an unknown d-dimensional mixed quantum state ρ, given few copies. We show that O(d/ε) copies suffice to obtain an estimate ρ that satisfies ||ρ − ρ||F2 ≤ ε (with high probability). An immediate consequence is that O((ρ) · d/ε2) ≤ O(d2/ε2) copies suffice to obtain an ε-accurate estimate in the standard trace distance. This improves on the best known prior result of O(d3/ε2) copies for full tomography, and even on the best known prior result of O(d2log(d/ε)/ε2) copies for spectrum estimation. Our result is the first to show that nontrivial tomography can be obtained using a number of copies that is just linear in the dimension. Next, we generalize these results to show that one can perform efficient principal component analysis on ρ. Our main result is that O(k d/ε2) copies suffice to output a rank-k approximation ρ whose trace-distance error is at most ε more than that of the best rank-k approximator to ρ. This subsumes our above trace distance tomography result and generalizes it to the case when ρ is not guaranteed to be of low rank. A key part of the proof is the analogous generalization of our spectrum-learning results: we show that the largest k eigenvalues of ρ can be estimated to trace-distance error ε using O(k2/ε2) copies. In turn, this result relies on a new coupling theorem concerning the Robinson–Schensted–Knuth algorithm that should be of independent combinatorial interest.
XIRHZ4U5
journalArticle
P. Facchi, Z. Hradil, G. Krenn, S. Pascazio, J. Řeháček
2002 · Physical Review A
Saved 2025-12-07
2LA9JUBD
journalArticle
Dominik Koutný, Libor Motka, Zdenıfmmode \checke\else ě\fik Hradil, Jaroslav ıfmmode \checkR\else Ř\fieháıfmmode \checkc\else č\fiek, Luis L. Sánchez-Soto
2022 · Phys. Rev. A · American Physical Society
Saved 2025-12-07
3V7GF4EM
document
Adibvafa Fallahpour, Andrew Magnuson, Purav Gupta, Shihao Ma, Jack Naimer, Arnav Shah, Haonan Duan, Omar Ibrahim et al.
2025
Saved 2025-12-07
P7CJ7BKC
webpage
Michael Nielsen
2018
Saved 2025-12-07
695JBHJL
journalArticle
Haoran Liao, Derek S. Wang, Iskandar Sitdikov, Ciro Salcedo, Alireza Seif, Zlatko K. Minev
2024 · Nature Machine Intelligence · Springer Science and Business Media LLC
Saved 2025-12-07
4JYJFIJH
book
Project Euler Problem Set
2022
Saved 2025-12-07
WK6VXPUX
book
Calculus I
Paul Dawkins
2022
Saved 2025-12-07
3EIRF8CL
document
Holger Krekel
2025
Saved 2025-12-07
NLNYHQAT
document
Hailan Ma, Zhenhong Sun, Daoyi Dong, Dong Gong
2025
Saved 2025-12-07
JSDBZXVH
document
Ben Hylek
2025
Saved 2025-12-07
3DFY7XK5
document
Hanrui Wang, Pengyu Liu, Kevin Shao, Dantong Li, Jiaqi Gu, David Z. Pan, Yongshan Ding, Song Han
2023
Saved 2025-12-07
EN994F8R
document
Haoran Wei, Yaofeng Sun, Yukun Li
2025
Saved 2025-12-07
9AW32GLH
book
René Descartes
1637 · Jan Maire
Saved 2025-12-07
X2RLULAX
book
Fyodor Dostoevsky
1880
Saved 2025-12-07
X7G3D3EP
document
Tianyi Li, Mingda Chen, Bowei Guo, Zhiqiang Shen
2025
Saved 2025-12-07
689I4ESD
book
Robert Chassel
1990
Saved 2025-12-07
RGEJTEJR
document
Ava Pun, Kangle Deng, Ruixuan Liu, Deva Ramanan, Changliu Liu, Jun-Yan Zhu
2025
Saved 2025-12-07
UX589N96
document
Keiron O'Shea, Ryan Nash
2015
Saved 2025-12-07
LARBQA7H
journalArticle
Hailan Ma, Bo Qi, Ian R Petersen, Re-Bing Wu, Herschel Rabitz, Daoyi Dong
2025 · National Science Review · Oxford University Press (OUP)
Saved 2025-12-07
SFA5YKDZ
document
Karan Kendre
2025
Saved 2025-12-07
YI2T4LGN
document
Zhikang Wang
2022
Saved 2025-12-07
XDBUAFVU
document
Omar Shindi, Qi Yu, Parth Girdhar, Daoyi Dong
2023
Saved 2025-12-07
53B5EJMU
journalArticle
V. V. Sivak, A. Eickbusch, H. Liu, B. Royer, I. Tsioutsios, M. H. Devoret
2022 · Phys. Rev. X · American Physical Society
Saved 2025-12-07
XMCLVT2R
document
Jiahao Yao, Paul Köttering, Hans Gundlach, Lin Lin, Marin Bukov
2020
Saved 2025-12-07
WPWVZKRJ
journalArticle
Yuval Baum, Mirko Amico, Sean Howell, Michael Hush, Maggie Liuzzi, Pranav Mundada, Thomas Merkh, Andre R.R. Carvalho et al.
2021 · PRX Quantum · American Physical Society (APS)
Saved 2025-12-07
RWCZSIKM
document
Ariel Norambuena, Marios Mattheakis, Francisco J. González, Raúl Coto
2023
Saved 2025-12-07
QJ92MX32
document
Abolfazl Ramezanpour
2025
Saved 2025-12-07
Y92DN5N7
document
Jan Ole Ernst, Tim Franzmeyer, Aniket Chatterjee, Axel Kuhn
2025
Saved 2025-12-07
PDVMG5KG
document
Yang Chen, Shaoshu Li
2016
Saved 2025-12-07
YVNLEG6J
document
Hang Xu, Tailong Xiao, Jingzheng Huang, Jianping Fan, Guihua Zeng
2025
Saved 2025-12-07
ZEL875YT
document
Florian Marquardt, Annett Püttmann
2008
Saved 2025-12-07
567ICF45
preprint
Karan Kendre
2025
Saved 2025-12-07
Quantum noise fundamentally limits the utility of near-term quantum devices, making error mitigation essential for practical quantum computation. While traditional quantum error correction codes require substantial qubit overhead and complex syndrome decoding, we propose a machine learning approach that directly reconstructs clean quantum states from noisy density matrices without additional qubits. We formulate quantum noise reduction as a supervised learning problem using a convolutional neural network (CNN) autoencoder architecture with a novel fidelity-aware composite loss function. Our method is trained and evaluated on a comprehensive synthetic dataset of 10,000 density matrices derived from random 5-qubit quantum circuits, encompassing five noise types (depolarizing, amplitude damping, phase damping, bit-flip, and mixed noise) across four intensity levels (0.05-0.20). The CNN successfully reconstructs quantum states across all noise conditions, achieving an average fidelity improvement from 0.298 to 0.774 (Δ = 0.476). Notably, the model demonstrates superior performance on complex mixed noise scenarios and higher noise intensities, with mixed noise showing the highest corrected fidelity (0.807) and improvement (0.567). The approach effectively preserves both diagonal elements (populations) and off-diagonal elements (quantum coherences), making it suitable for entanglement-dependent quantum algorithms. While phase damping presents fundamental information-theoretic limitations, our results suggest that CNN-based density matrix reconstruction offers a promising, resource-efficient alternative to traditional quantum error correction for NISQ-era devices. This data-driven approach could enable practical quantum advantage with fewer physical qubits than conventional error correction schemes require.
QD8KGV3K
preprint
M. Mattheakis, P. Protopapas, D. Sondak, M. Di Giovanni, E. Kaxiras
2020
Saved 2025-12-07
Neural networks are a central technique in machine learning. Recent years have seen a wave of interest in applying neural networks to physical systems for which the governing dynamics are known and expressed through differential equations. Two fundamental challenges facing the development of neural networks in physics applications is their lack of interpretability and their physics-agnostic design. The focus of the present work is to embed physical constraints into the structure of the neural network to address the second fundamental challenge. By constraining tunable parameters (such as weights and biases) and adding special layers to the network, the desired constraints are guaranteed to be satisfied without the need for explicit regularization terms. This is demonstrated on upervised and unsupervised networks for two basic symmetries: even/odd symmetry of a function and energy conservation. In the supervised case, the network with embedded constraints is shown to perform well on regression problems while simultaneously obeying the desired constraints whereas a traditional network fits the data but violates the underlying constraints. Finally, a new unsupervised neural network is proposed that guarantees energy conservation through an embedded symplectic structure. The symplectic neural network is used to solve a system of energy-conserving differential equations and out-performs an unsupervised, non-symplectic neural network.