Zotero Catalogue

A public reading list of papers, books, videos, and other resources. The inclusion of a resource on this catalogue is NOT an endorsement of anything contained within, and in most cases the resources has not been read by me at the time of saving.

1157 items · showing 801–850 · page 17 of 24 Sort: Newest Oldest Title A–Z Title Z–A

36TWY7KW
webpage
Saved 2026-05-06
FFZE8LCK
journalArticle
Manifesto of the Communist Party
Karl Marx
Saved 2026-05-04
CIZWPSMD
preprint
George Morgulis, John Hewitt
2026
Saved 2026-05-04
Subliminal learning describes a student language model inheriting a behavioral bias by fine-tuning on seemingly innocuous data generated by a biased teacher model. Prior work has begun to characterize this phenomenon but leaves open questions about the scope of signals it can transfer, the mechanisms that explain it, and the precision with which a bias can be encoded by seemingly unrelated data. We tackle all three problems by introducing subliminal steering, a variant of subliminal learning in which the teacher’s bias is implemented not via a system prompt, as in prior work, but through a steering vector trained to maximize the likelihood of a set of target samples. First, we show that subliminal steering transfers complex multi-word biases, whereas prior work focused on single-word preferences—demonstrating a large scope of subliminally transferrable signals. Second, we provide mechanistic evidence that subliminal learning transfers not only the target behavioral bias, but also the steering vector itself, localized to the layers at which the teacher was steered. Finally, we show that the bias is encoded with surprising precision. We train a new steering vector directly on the subliminally-laden dataset and find that it attains high cosine similarity with the original vector.
RGXBTLDN
webpage
Elon Litman
2025
Saved 2026-05-04
How the search for a simple approximation revealed a beautiful, exact solution.
DH5ZUCEJ
blogPost
Laura Schroeder
2026
Saved 2026-05-04
on clarity, endurance, and putting yourself out there
LLP359S6
blogPost
Tea
2010
Saved 2026-05-03
AF7AF4WW
preprint
Thomas Hayes, Roshan Rao, Halil Akin, Nicholas J. Sofroniew, Deniz Oktay, Zeming Lin, Robert Verkuil, Vincent Q. Tran et al.
2024
Saved 2026-05-02
More than three billion years of evolution have produced an image of biology encoded into the space of natural proteins. Here we show that language models trained on tokens generated by evolution can act as evolutionary simulators to generate functional proteins that are far away from known proteins. We present ESM3, a frontier multimodal generative language model that reasons over the sequence, structure, and function of proteins. ESM3 can follow complex prompts combining its modalities and is highly responsive to biological alignment. We have prompted ESM3 to generate fluorescent proteins with a chain of thought. Among the generations that we synthesized, we found a bright fluorescent protein at far distance (58% identity) from known fluorescent proteins. Similarly distant natural fluorescent proteins are separated by over five hundred million years of evolution.
LRKKSZSU
forumPost
Yacine Mahdid [@yacinelearning]
2025
Saved 2026-05-01
P7XNVSX2
blogPost
Derrick Harris, Matt Bornstein, Guido Appenzeller
2023
Saved 2026-05-01
A curated list of resources we’ve relied on to get smarter about modern AI, including generative AI, LLMs, and transformer models.
DAM7J2SN
webpage
Saved 2026-04-29
XKRN9D3D
webpage
Ross Benes
Saved 2026-04-28
XI3AXMD9
magazineArticle
· The Economist
Saved 2026-04-28
L5ENGT2K
webpage
Tim Richardson
Saved 2026-04-28
'Severe punishment' for online miscreants
6XAW6L2T
blogPost
Tyler Cowen
2026
Saved 2026-04-28
From MR commentator Sure: Generally such figures do not reside within the physicians’ office. On our side of the table we do some procedure with multiple specifications and generate some CPT code(s) (e.g. a lap cholycystectomy is 47562, add on a common bile duct exploration and it becomes a 47564, and if you just do […]
UUDZIXDS
webpage
Kelsey Piper
2026
Saved 2026-04-28
AI only needs 150 words to identify me. What does that mean for you?
ZCECW6F6
webpage
2026
Saved 2026-04-28
AI can echolocate authors through their prose. Your digital fingerprint is at risk.
7U75TIEQ
webpage
Saved 2026-04-28
We create innovative outdoor products and experiences that fund sustainable poverty relief, move people to do good, and inspire adventure.
LP4T5FPQ
webpage
Saved 2026-04-28
Andela provides the human compute layer behind modern AI systems — training models,
 deploying AI-native engineers, and upskilling the teams that build them.
GT6Q8XAR
blogPost
Aaryan Harshith
2026
Saved 2026-04-27
And why reading it is the most meaningful challenge in biology
GZLB93BP
journalArticle
Semantic Self-Consistency: Enhancing Language Model Reasoning via Semantic Weighting
Tim Knappe, Ryan Li, Ayush Chauhan, Kaylee Chhua, Kevin Zhu, Sean O’Brien
Saved 2026-04-26
While large language models (LLMs) have rapidly improved their performance on a broad number of tasks, they still often fall short on reasoning tasks. As LLMs become more integrated in diverse real-world tasks, advancing their reasoning capabilities is crucial to their effectiveness in nuanced, complex problems. Wang et al. [33]’s self-consistency framework reveals that sampling multiple rationales before taking a majority vote reliably improves model performance across various closed-answer reasoning tasks. Standard methods based on this framework aggregate the final decisions of these rationales but fail to utilize the semantic information detailed in the step-by-step reasoning paths. Our work introduces semantic self-consistency, enhancing this approach by incorporating and analyzing both the reasoning paths of these rationales in addition to their final decisions before taking a majority vote. These methods not only improve the reliability of reasoning paths but also cause more robust performance on complex reasoning tasks.
VAMTKFWW
preprint
Timothy P. Lillicrap, Daniel Cownden, Douglas B. Tweed, Colin J. Akerman
2014
Saved 2026-04-26
The brain processes information through many layers of neurons. This deep architecture is representationally powerful, but it complicates learning by making it hard to identify the responsible neurons when a mistake is made. In machine learning, the backpropagation algorithm assigns blame to a neuron by computing exactly how it contributed to an error. To do this, it multiplies error signals by matrices consisting of all the synaptic weights on the neuron's axon and farther downstream. This operation requires a precisely choreographed transport of synaptic weight information, which is thought to be impossible in the brain. Here we present a surprisingly simple algorithm for deep learning, which assigns blame by multiplying error signals by random synaptic weights. We show that a network can learn to extract useful information from signals sent through these random feedback connections. In essence, the network learns to learn. We demonstrate that this new mechanism performs as quickly and accurately as backpropagation on a variety of problems and describe the principles which underlie its function. Our demonstration provides a plausible basis for how a neuron can be adapted using error signals generated at distal locations in the brain, and thus dispels long-held assumptions about the algorithmic constraints on learning in neural circuits.
RKVKZ7LK
blogPost
David Stutz
2024
Saved 2026-04-26
The decision to have a separate High School Project Track at NeurIPS 2024 has sparked quite some controversy, with many prominent AI researchers debating pros and cons and personal opinions, primarily on X/Twitter. Initially, I ignored this discussion, but eventually started thinking about it myself. Here are some of my thoughts.
IY87HZXJ
webpage
Zilan Qian
2025
Saved 2026-04-25
Life inside the "human sea attack"
53LZVIDP
journalArticle
Maison Clouâtré, Stefano Marano, Peter L. Falb, Moe Z. Win
2024 · IEEE Control Systems Letters
Saved 2026-04-24
This letter investigates parameter estimation in quantum systems that undergo dynamical evolution. Optimal control problems are formulated to maximize the information, about an unknown parameter, extracted by a given quantum measurement apparatus. This letter introduces the concept of “admissible controls”—control laws that do not depend on the unknown parameter they elicit. For scalar parameter estimation in unital quantum systems interrogated by binary measurements, this letter derives a necessary and sufficient condition on quantum measurement operators so that an information maximizing control law is admissible. When the admissibility condition is satisfied, it is shown that the resulting optimal control problem may be solved using well-established techniques.
BBLRCC9C
blogPost
Chamath Palihapitiya
2026
Saved 2026-04-24
2025 was, in some ways, a historic year for Social Capital. Numerically, we did well. Our portfolio took an important inflection upwards when NVIDIA licensed Groq for $20B. However...
3J65XLZ8
webpage
Erik Torenberg
2026
Saved 2026-04-24
America | Tech | Opinion | Culture | Charts
4CWU4XEN
webpage
Erik Torenberg
2026
Saved 2026-04-24
America | Tech | Opinion | Culture | Charts
TH4A97RU
journalArticle
John A. Smolin, Jay M. Gambetta, Graeme Smith
2012 · Physical Review Letters
Saved 2026-04-24
We provide an efficient method for computing the maximum likelihood mixed quantum state (with density matrix $ρ$) given a set of measurement outcome in a complete orthonormal operator basis subject to Gaussian noise. Our method works by first changing basis yielding a candidate density matrix $μ$ which may have nonphysical (negative) eigenvalues, and then finding the nearest physical state under the 2-norm. Our algorithm takes at worst $O(d^4)$ for the basis change plus $O(d^3)$ for finding $ρ$ where $d$ is the dimension of the quantum state. In the special case where the measurement basis is strings of Pauli operators, the basis change takes only $O(d^3)$ as well. The workhorse of the algorithm is a new linear-time method for finding the closest probability distribution (in Euclidean distance) to a set of real numbers summing to one.
329SGRQG
webpage
2026
Saved 2026-04-23
Introducing GPT-5.5, our smartest model yet—faster, more capable, and built for complex tasks like coding, research, and data analysis across tools.
UEXE69RH
webpage
2026
Saved 2026-04-23
We’ve redesigned our engineering interview process from the ground up.
Q2VCE2FT
blogPost
Chamath Palihapitiya
2026
Saved 2026-04-22
2025 was, in some ways, a historic year for Social Capital. Numerically, we did well. Our portfolio took an important inflection upwards when NVIDIA licensed Groq for $20B. However...
T3QPJUZT
blogPost
Tyler Cowen
2019
Saved 2026-04-22
Following up on my post a few days ago, about the value of deliberate practice for knowledge workers, a number of you asked me what form my practice takes.  A few of you were skeptical, but it is long since established that practice improves both your writing and your memory, so surely it can do […]
I4EH752N
webpage
Substack
Saved 2026-04-21
why the most respected engineers are pushing back on the coding agent hype cycle, and what the claude code argument actually reveals
KMRGEYTT
book
Quantum circuit fidelity estimation using machine learning
Avi Vadali, Rutuja Kshirsagar, Prasanth Shyamsundar, Gabriel Perdue
2022
Saved 2026-04-21
The computational power of real-world quantum computers is limited by errors. When using quantum computers to perform algorithms which cannot be efficiently simulated classically, it is important to quantify the accuracy with which the computation has been performed. In this work we introduce a machine-learning-based technique to estimate the fidelity between the state produced by a noisy quantum circuit and the target state corresponding to ideal noise-free computation. Our machine learning model is trained in a supervised manner, using smaller or simpler circuits for which the fidelity can be estimated using other techniques like direct fidelity estimation and quantum state tomography. We demonstrate that the trained model can predict the fidelities of more complicated circuits for which such methods are infeasible.
JRRUIFID
webpage
Saved 2026-04-21
Why Xanadu’s hype-averse QML team is betting on the quantum Fourier transform to advance machine learning and better machine learning models.
87RKDS9E
blogPost
Alex Tabarrok
2026
Saved 2026-04-21
It’s well known that grade inflation has “degraded” the informational content of grades at many colleges. At Harvard, two-thirds of all undergraduate grades are now A’s—up from about a quarter two decades ago. In response, a Harvard faculty committee has proposed capping A grades at 20 percent of each class (plus a cushion for small […]
LE77V83K
webpage
Tony Kulesa
2023
Saved 2026-04-21
Tyler Cowen is an economist that has spotted top talent in fields ranging from biotech to literature, often years before insiders. How does he do it?
VUZVQQTY
webpage
Substack
Saved 2026-04-21
addiction is good, actually.
XEEEG4DZ
webpage
Substack
Saved 2026-04-21
What will happen in our second peasanthood
NAHF6KQH
book
American Prometheus: The Triumph and Tragedy of J. Robert Oppenheimer
Kai Bird, Martin Sherwin
Saved 2026-04-20
QZAF8EG4
webpage
Craig Mod
Saved 2026-04-20
Essay on the benefits of speedy software, and how it affects user perception of engineering quality and overall usability
9R5JZPPD
webpage
Substack
Saved 2026-04-19
that's the artist's vocation for ya
9WU9WEJE
webpage
Saved 2026-04-19
Spawn coding agents that run infinitely in the cloud. Powered by AI SDK, Gateway, Sandbox, and Workflow SDK.
TJ4WXXPD
preprint
Christos Louizos, Max Welling, Diederik P. Kingma
2018
Saved 2026-04-19
We propose a practical method for $L_0$ norm regularization for neural networks: pruning the network during training by encouraging weights to become exactly zero. Such regularization is interesting since (1) it can greatly speed up training and inference, and (2) it can improve generalization. AIC and BIC, well-known model selection criteria, are special cases of $L_0$ regularization. However, since the $L_0$ norm of weights is non-differentiable, we cannot incorporate it directly as a regularization term in the objective function. We propose a solution through the inclusion of a collection of non-negative stochastic gates, which collectively determine which weights to set to zero. We show that, somewhat surprisingly, for certain distributions over the gates, the expected $L_0$ norm of the resulting gated weights is differentiable with respect to the distribution parameters. We further propose the \emph{hard concrete} distribution for the gates, which is obtained by "stretching" a binary concrete distribution and then transforming its samples with a hard-sigmoid. The parameters of the distribution over the gates can then be jointly optimized with the original network parameters. As a result our method allows for straightforward and efficient learning of model structures with stochastic gradient descent and allows for conditional computation in a principled way. We perform various experiments to demonstrate the effectiveness of the resulting approach and regularizer.
KT5885AI
preprint
Jiacheng Liu, Xiaohan Zhao, Xinyi Shang, Zhiqiang Shen
2026
Saved 2026-04-18
Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study describes its comprehensive architecture by analyzing the publicly available TypeScript source code and further comparing it with OpenClaw, an independent open-source AI agent system that answers many of the same design questions from a different deployment context. Our analysis identifies five human values, philosophies, and needs that motivate the architecture (human decision authority, safety and security, reliable execution, capability amplification, and contextual adaptability) and traces them through thirteen design principles to specific implementation choices. The core of the system is a simple while-loop that calls the model, runs tools, and repeats. Most of the code, however, lives in the systems around this loop: a permission system with seven modes and an ML-based classifier, a five-layer compaction pipeline for context management, four extensibility mechanisms (MCP, plugins, skills, and hooks), a subagent delegation mechanism with worktree isolation, and append-oriented session storage. A comparison with OpenClaw, a multi-channel personal assistant gateway, shows that the same recurring design questions produce different architectural answers when the deployment context changes: from per-action safety classification to perimeter-level access control, from a single CLI loop to an embedded runtime within a gateway control plane, and from context-window extensions to gateway-wide capability registration. We finally identify six open design directions for future agent systems, grounded in recent empirical, architectural, and policy literature.