Zotero Catalogue

A public reading list of papers, books, videos, and other resources. The inclusion of a resource on this catalogue is NOT an endorsement of anything contained within, and in most cases the resources has not been read by me at the time of saving.

1157 items · showing 551–600 · page 12 of 24 Sort: Newest Oldest Title A–Z Title Z–A

JAXNTJIR
preprint
Ofir Press, Noah A. Smith, Mike Lewis
2022
Saved 2026-06-21
Since the introduction of the transformer model by Vaswani et al. (2017), a fundamental question has yet to be answered: how does a model achieve extrapolation at inference time for sequences that are longer than it saw during training? We first show that extrapolation can be enabled by simply changing the position representation method, though we find that current methods do not allow for efficient extrapolation. We therefore introduce a simpler and more efficient position method, Attention with Linear Biases (ALiBi). ALiBi does not add positional embeddings to word embeddings; instead, it biases query-key attention scores with a penalty that is proportional to their distance. We show that this method trains a 1.3 billion parameter model on input sequences of length 1024 that extrapolates to input sequences of length 2048, achieving the same perplexity as a sinusoidal position embedding model trained on inputs of length 2048 but training 11% faster and using 11% less memory. ALiBi's inductive bias towards recency also leads it to outperform multiple strong position methods on the WikiText-103 benchmark.
QJTRMLQ6
journalArticle
Hailan Ma, Zhenhong Sun, Daoyi Dong, Dong Gong
2026 · IEEE Transactions on Emerging Topics in Computational Intelligence
Saved 2026-06-21
Quantum state tomography (QST) is the process of reconstructing the complete state of a quantum system (mathematically described as a density matrix) through a series of different measurements. These measurements are performed on a number of identical copies of the quantum system, with outcomes gathered as probabilities/frequencies. QST aims to recover the density matrix and the corresponding properties of the quantum state from the measured frequencies. Although an informationally complete set of measurements can specify the quantum state accurately in an ideal scenario with a large number of identical copies, both the measurements and identical copies are restricted and imperfect in practical scenarios, making QST highly ill-posed. The conventional QST methods usually assume adequate or accurate measured frequencies or rely on manually designed regularizers to handle the ill-posed reconstruction problem, suffering from limited applications in realistic scenarios. Recent advances in deep neural networks (DNNs) led to the emergence of deep learning (DL) in QST. However, existing DL-based QST approaches often employ generic DNN models that are not optimized for imperfect conditions of QST. In this paper, we propose a transformerbased autoencoder architecture tailored for QST with imperfect measurement data. Our method leverages a transformer-based encoder to extract an informative latent representation (ILR) from imperfect measurement data and employs a decoder to predict the quantum states based on the ILR. We anticipate that the high-dimensional ILR will capture more comprehensive information about the quantum states. To achieve this, we conduct pre-training of the encoder using a pretext task that involves reconstructing high-quality frequencies from measured frequencies. Extensive simulations and experiments demonstrate the remarkable ability of the informative latent representation to deal with imperfect measurement data in QST.
PJ7M2ZUX
forumPost
Zvi
2023
Saved 2026-06-20
BJKMM4UB
blogPost
Zvi Mowshowitz
2022
Saved 2026-06-20
Or: Against Car Seat Laws At Least Beyond Age 2
QVB66AFX
webpage
Zilan Qian
2026
Saved 2026-06-19
and what it all means
ED6ES2UQ
webpage
Farhan Thawar
2026
Saved 2026-06-19
Introduce the Canadian AI Subscription Deduction (CASD) for individuals: a 100 percent personal tax deduction of up to $3,000 per year for qualifying AI subscriptions, learning tools, and productivity services.
7TWAXYTE
webpage
Saved 2026-06-18
A New Era of Midjourney, announcing Midjourney Medical: full-body Ultrasonic CT and the Midjourney Spa.
L6TCG85B
webpage
Saved 2026-06-18
Capture every rollout, score what happened, and retrain on the parts that matter.
4DEG8XK2
preprint
Amil Dravid, Yasaman Bahri, Alexei A. Efros, Yossi Gandelsman
2026
Saved 2026-06-18
We investigate whether neuron populations within neural networks evolve predictably with scale, extending scaling laws beyond macroscopic observables such as loss. To probe this question, we study Rosetta Neurons, a previously characterized class of neurons whose activation patterns are similar across independently trained models (Dravid et al., 2023). In separate analyses of language models up to 30B parameters and vision models up to 5B parameters, we observe that the population of Rosetta Neurons follows a sublinear power law in model size, growing in absolute number but occupying a shrinking fraction of the total neuron count. We further observe a Neuron Polarization Effect: Rosetta Neurons become more selective and increasingly monosemantic with scale, separating from a growing non-Rosetta population that remains less selective. An analytical model balancing feature utility against limited neuron capacity explains the sublinear power-law scaling and this polarization effect. Finally, we find that Rosetta Neurons become more domain-specialized with scale and illustrate their selectivity through a targeted data-filtering case study for continued pretraining. Our results point to a scaling law for interpretable, shared neuron-level structure, linking model size to systematic changes in neuron universality, selectivity, and specialization.
55J2BQMI
webpage
2026
Saved 2026-06-18
A Field Guide to AI Fellowships If you’re early in AI and you keep hearing “you should apply to a fellowship” without anyone telling you which one, for what, or how — this is for you.
NI2CCF23
forumPost
vivek [@itsreallyvivek]
2026
Saved 2026-06-18
JVXM69UA
forumPost
Elizabeth
2025
Saved 2026-06-17
Y6CDCR8G
journalArticle
Gwern
2026
Saved 2026-06-17
How much to edit? A top-𝑘 attention-window toy model of the publish-polish trade off: You should polish in inverse proportion to how much of your writing a reader will see; the more diverse your writing, the better off you are writing 𝑚𝑜𝑟𝑒 rather than 𝑏𝑒𝑡𝑡𝑒𝑟.
FTGG63SB
journalArticle
Monitoring Ultrafast Lattice Dynamics in 2D NbTe2
Christian Viernes
Saved 2026-06-17
QZE7IBUP
journalArticle
Eoin Whelan
2026 · Cyberpsychology, Behavior, and Social Networking
Saved 2026-06-17
The issue of whether social media use does or does not influence adolescent well-being remains a pressing concern for policymakers, parents, and researchers. As evidence of the harmful effects of social media, some point to the fact that young adults frequently regret the time they spent on social media when they were younger. However, highlighting a single instance of retrospective regret as a proxy for harm may be misleading. Our objective in this study is to provide a more accurate and contextually grounded assessment of social media regrets by benchmarking them against a broader set of common teenage regrets, as rated by young adults. We then determine which teenage regrets predict current life satisfaction. Four hundred young adults aged 20–24 from Ireland, the United Kingdom, and the United States were recruited via the Prolific platform and completed an online survey assessing the potency of 20 retrospective regrets when they reflect on their teenage years. Our findings indicate that social media was not a prominent source of regret when young adults reflect on their teenage years. In addition, regrets about social media use were not associated with current life satisfaction, but other regrets were. As such, our findings are consistent with other prior studies suggesting that at the population level, the harmful effects of social media use on adolescent well-being may be overstated. However, these results should be interpreted with caution and may not reflect the experience of all individuals within this population.
9BWCMH9H
preprint
Michael Broughton, Guillaume Verdon, Trevor McCourt, Antonio J. Martinez, Jae Hyeon Yoo, Sergei V. Isakov, Philip Massey, Ramin Halavati et al.
2021
Saved 2026-06-16
We introduce TensorFlow Quantum (TFQ), an open source library for the rapid prototyping of hybrid quantum-classical models for classical or quantum data. This framework offers high-level abstractions for the design and training of both discriminative and generative quantum models under TensorFlow and supports high-performance quantum circuit simulators. We provide an overview of the software architecture and building blocks through several examples and review the theory of hybrid quantum-classical neural networks. We illustrate TFQ functionalities via several basic applications including supervised learning for quantum classification, quantum control, simulating noisy quantum circuits, and quantum approximate optimization. Moreover, we demonstrate how one can apply TFQ to tackle advanced quantum learning tasks including meta-learning, layerwise learning, Hamiltonian learning, sampling thermal states, variational quantum eigensolvers, classification of quantum phase transitions, generative adversarial networks, and reinforcement learning. We hope this framework provides the necessary tools for the quantum computing and machine learning research communities to explore models of both natural and artificial quantum systems, and ultimately discover new quantum algorithms which could potentially yield a quantum advantage.
CQSE2IRX
webpage
2009
Saved 2026-06-16
H7UTFHSY
blogPost
Raghuveer Parthasarathy
2026
Saved 2026-06-16
What makes a college course popular or unpopular? I’ve long been interested in courses for non-science majors that satisfy “general education” requirements, their aim being to fos…
DBIIQDYN
webpage
Saved 2026-06-16
97LA4BML
preprint
Bruno Gavranović, Paul Lessard, Andrew Dudzik, Tamara von Glehn, João G. M. Araújo, Petar Veličković
2024
Saved 2026-06-16
We present our position on the elusive quest for a general-purpose framework for specifying and studying deep learning architectures. Our opinion is that the key attempts made so far lack a coherent bridge between specifying constraints which models must satisfy and specifying their implementations. Focusing on building a such a bridge, we propose to apply category theory -- precisely, the universal algebra of monads valued in a 2-category of parametric maps -- as a single theory elegantly subsuming both of these flavours of neural network design. To defend our position, we show how this theory recovers constraints induced by geometric deep learning, as well as implementations of many architectures drawn from the diverse landscape of neural networks, such as RNNs. We also illustrate how the theory naturally encodes many standard constructs in computer science and automata theory.
8R675IYK
blogPost
Joshua Gans
2026
Saved 2026-06-16
Straight out of Magic
V59CHHQ3
bookSection
Tanmay Sah, Vishal Srivastava, Dolly Sah, Kayden Jordan
2026 · Proceedings of the ACM Conference on AI and Agentic Systems · Association for Computing Machinery
Saved 2026-06-14
We study how runtime enforcement against unsafe actions affects end-to-end task performance in multi-step tool using large language model (LLM) agents. Using τ -bench across Airline and Retail domains, we compare baseline Tool-Calling, planning-integrated (Triad), and policy-mediated (Triad-Safety) architectures with GPT-OSS-20B and GLM-4-9B. We identify model dependent interaction horizons (15–30 turns) and decompose outcomes into overall success rate (SR), safe success rate (SSR), and unsafe success rate (USR). Our results reveal a persistent “Safety-Capability Gap”. While safety mediation can intercept up to 94% of non-compliant actions, it rarely translates into strictly safe goal attainment (SSR < 5% in most settings). We find that high unsafe success rates are primarily driven by “Integrity Leaks,” where models hallucinate user identifiers to bypass mandatory authentication. Recovery rates following blocked actions are consistently low, ranging from 21% for GPT-OSS-20B in simpler procedural tasks to near 0% in complex Retail scenarios. These results demonstrate that runtime enforcement imposes a significant “verifier tax” on conversational length and compute cost without guaranteeing safe completion, highlighting the critical need for agents capable of grounded identity verification and post-intervention reasoning.
97UXFNH4
blogPost
Tyler Cowen
2022
Saved 2026-06-14
Teenage mental health has been a source of growing concern over the past decade, with recent whistleblower testimony pointing to the mental health risks of spending time on social media platforms, especially for girls. This paper investigates the extent to which social media are harmful for teenagers, leveraging rich administrative data from the Canadian province […]
WE59LKIY
preprint
Yarin Gal, Zoubin Ghahramani
2016
Saved 2026-06-13
Deep learning tools have gained tremendous attention in applied machine learning. However such tools for regression and classification do not capture model uncertainty. In comparison, Bayesian models offer a mathematically grounded framework to reason about model uncertainty, but usually come with a prohibitive computational cost. In this paper we develop a new theoretical framework casting dropout training in deep neural networks (NNs) as approximate Bayesian inference in deep Gaussian processes. A direct result of this theory gives us tools to model uncertainty with dropout NNs -- extracting information from existing models that has been thrown away so far. This mitigates the problem of representing uncertainty in deep learning without sacrificing either computational complexity or test accuracy. We perform an extensive study of the properties of dropout's uncertainty. Various network architectures and non-linearities are assessed on tasks of regression and classification, using MNIST as an example. We show a considerable improvement in predictive log-likelihood and RMSE compared to existing state-of-the-art methods, and finish by using dropout's uncertainty in deep reinforcement learning.
YQ3F63JJ
webpage
Saved 2026-06-12
What happens in a world where AIs make scientific discoveries that humans cannot understand?
G8YUKWHH
conferencePaper
Susumu Kuno, Anthony G. Oettinger
1963 · Proceedings of the November 12-14, 1963, fall joint computer conference on XX - AFIPS '63 (Fall) · ACM Press
Saved 2026-06-12
Z5492T9W
webpage
Saved 2026-06-12
YTLL8HPM
webpage
Daan Juijn, Stan van Baarsen, Judith Dada, Lily Stelling, Philip Fox, Alex Petropoulos, Michiel Bakker
Saved 2026-06-12
A five-year scenario about AI and Europe's impending slide into irrelevance, with a 2034 epilogue that describes how the collapse of the European model could have been prevented.
JFTBBIGW
webpage
Saved 2026-06-12
Learn faster, achieve more
3E3URGC8
webpage
Saved 2026-06-12
Jeri Ellsworth documents her amateur science experiments.
6D9S2D29
blogPost
2018
Saved 2026-06-12
FUMMYSYE
book
Functions 11
Saved 2026-06-11
RFXY7AYC
webpage
2026
Saved 2026-06-11
By Ryan Lopopolo, Member of the Technical Staff
U7NY6RHF
webpage
2026
Saved 2026-06-11
A case against measuring AI-assisted engineering by code volume and other vanity metrics instead of outcomes.
VPLW9QN9
journalArticle
How Does Batch Normalization Help Optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, Aleksander Ma
Saved 2026-06-11
Batch Normalization (BatchNorm) is a widely adopted technique that enables faster and more stable training of deep neural networks (DNNs). Despite its pervasiveness, the exact reasons for BatchNorm’s effectiveness are still poorly understood. The popular belief is that this effectiveness stems from controlling the change of the layers’ input distributions during training to reduce the so-called “internal covariate shift”. In this work, we demonstrate that such distributional stability of layer inputs has little to do with the success of BatchNorm. Instead, we uncover a more fundamental impact of BatchNorm on the training process: it makes the optimization landscape significantly smoother. This smoothness induces a more predictive and stable behavior of the gradients, allowing for faster training.
5KL6GC7U
forumPost
luuk0987
2018
Saved 2026-06-11
MMQCV9WV
webpage
Saved 2026-06-10
Repository for the book "Crafting Interpreters". Contribute to munificent/craftinginterpreters development by creating an account on GitHub.
XAZAR9WW
webpage
Saved 2026-06-10
3VWBJEJ8
webpage
Saved 2026-06-10
M7B78Y4F
journalArticle
NBER WORKING PAPER SERIES
Caitlin K Myers, Ezekiel Hooper
Saved 2026-06-09
The U.S. general fertility rate has fallen by 22% since 2007, a sustained decline not readily explained by economic conditions, contraceptive use, housing or childcare costs, or other commonly cited factors. We assess the potential role of a different shock: the diffusion of the smartphone. The U.S. rollout of the iPhone, the first modern smartphone, provides a natural experiment: from June 2007 through February 2011, the device was sold only on AT&T, allowing us to identify its effect from variation in AT&T’s mobile broadband coverage. Entropy-balanced Poisson and synthetic difference-in-differences event studies imply that access to the iPhone reduced births by 4.5–8.0% at ages 15–19 and 3.2–6.6%at ages 20–24, with statistically significant but smaller declines among older cohorts. Placebo analyses applied to Verizon and Sprint’s pre-2011 coverage footprint are null. Taken together, these cohort effects imply that the diffusion of the iPhone deepened the decline in births among women under 30 while suppressing the rise in births among older women. Overall, the diffusion of the iPhone explains 33–52% of the decline in the general fertility rate among women aged 15–44. National-survey evidence on time use and sexual behavior is consistent with the iPhone reducing in-person interactions, increasing pornography use, and reducing sexual frequency.
HC6KGQNV
blogPost
Tyler Cowen
2026
Saved 2026-06-09
The U.S. general fertility rate has fallen by 22% since 2007, a sustained decline not readily explained by economic conditions, contraceptive use, housing or childcare costs, or other commonly cited factors. We assess the potential role of a different shock: the diffusion of the smartphone. The U.S. rollout of the iPhone, the first modern smartphone, […]
GZE4BQ5V
webpage
german s
2026
Saved 2026-06-08
The act of pumping immense, disproportionate resources into a previously casual or complex and layered activity to forcefully extract and squeeze out the purest, most concentrated dopamine hit
E6JBKC69
webpage
Saved 2026-06-08
Five exercises for building research taste (and three failure modes).
54GF6Z42
webpage
Saved 2026-06-07
5KPHIB45
journalArticle
Clement Gosselin, Thierry Laliberte, Audrey Veillette
2015 · IEEE Transactions on Robotics
Saved 2026-06-06