Zotero Catalogue
A public reading list of papers, books, videos, and other resources. The inclusion of a resource on this catalogue is NOT an endorsement of anything contained within, and in most cases the resources has not been read by me at the time of saving.
1157 items · showing 501–550 · page 11 of 24 Sort: Newest Oldest Title A–Z Title Z–A
MUSZDCUL
blogPost
Danielle Strachman,
areoform
1517's Teen Fall Camp is back! So back!
536ERY36
webpage
Kaitlyn Tiffany
An evolutionary psychologist is challenging the popular understanding of kids and technology.
X9ZN5SRB
webpage
Jean M. Twenge
New analyses from Europe, Japan, and South Korea show increases that in many cases began before COVID.
7HKSREFF
blogPost
Michael Roberts
Part One: the revolution On the 4th July 2026 it will be 250 years since the 13 British colonies in North America declared independence from the British colonial power at a congress in Philadelphia…
HDL8X76X
webpage
A detailed forecast and recommendation for how the US, China and the rest of the world should navigate superintelligence.
FB6UIW5W
blogPost
Alex Tabarrok
Spider Noir (Prime): I’ve had enough of the Marvel multiverse so I was worried about Spider-Noir. The writers, however, have written an excellent noir in the style of Raymond Chandler with Nicholas Cage channeling Humphrey Bogart. The Spiderman stuff is all there but it is appropriately embedded. There are some excellent lines. Most notably an […]
BM5PRQC7
webpage
thinking about scary things • examples from Wave • examples from elsewhere • finding a buddy • getting the timing right • a list of abyss questions
9FEZR4TE
journalArticle
Notes on Diffy Qs: Differential Equations for Engineers
Jiri Lebl
M57VR6M2
webpage
Need a VectorDB for Your GenAI Apps?Zilliz Cloud is a managed vector database built on Milvus perfect for building GenAI applications Try Free
Mean pooling is commonly used to create sentence embeddings from transformer token outputs like BERT because it provides
3ZC4AYWE
webpage
Brian Potter
Electric power in the US is provided by the electrical grid, a huge network of power plants, transmission lines, and transformers that moves electric power from where it's generated to where it's consumed.
225YTHCC
preprint
Zhe Yin,
Xiaodong Gu,
Beijun Shen
Code language models excel on code intelligence tasks, yet their internal interpretability is underexplored. Existing neuron interpretability techniques from NLP are suboptimal for source code due to programming languages formal, hierarchical, and executable nature. We empirically investigate code LLMs at the neuron level, localizing language-specific neurons (selectively responsive to one language) and concept layers (feed-forward layers encoding language-agnostic code representations). We analyze Llama-3.1-8B and Qwen2.5-Coder-32B on multilingual inputs in C++, Java, Python, Go, and JavaScript, measuring neuron selectivity and layerwise contributions during generation. We find (1) neurons specialized for individual languages alongside a universal subset supporting general-purpose generation; and (2) lower layers mainly encode language-specific syntax, while middle layers capture semantic abstractions shared across languages, emerging as concept layers. We demonstrate utility on three tasks: neuron-guided fine-tuning for code generation, clone detection via concept-layer embeddings, and concept-layer-guided transfer for code summarization, each yielding consistent gains in multilingual settings.
DD4WMCVR
journalArticle
Guruprasad Raghavan,
Bahey Tharwat,
Surya Narayanan Hari,
Dhruvil Satani,
Rex Liu,
Matt Thomson
Contemporary machine learning algorithms train artificial neural networks by setting network weights to a single optimized configuration through gradient descent on task-specific training data. The resulting networks can achieve human-level performance on natural language processing, image analysis and agent-based tasks, but lack the flexibility and robustness characteristic of human intelligence. Here we introduce a differential geometry framework—functionally invariant paths—that provides flexible and continuous adaptation of trained neural networks so that secondary tasks can be achieved beyond the main machine learning goal, including increased network sparsification and adversarial robustness. We formulate the weight space of a neural network as a curved Riemannian manifold equipped with a metric tensor whose spectrum defines low-rank subspaces in weight space that accommodate network adaptation without loss of prior knowledge. We formalize adaptation as movement along a geodesic path in weight space while searching for networks that accommodate secondary objectives. With modest computational resources, the functionally invariant path algorithm achieves performance comparable with or exceeding state-of-the-art methods including low-rank adaptation on continual learning, sparsification and adversarial robustness tasks for large language models (bidirectional encoder representations from transformers), vision transformers (ViT and DeIT) and convolutional neural networks.
RLS6SVT4
journalArticle
Exploring non-invasive sexing of early chick embryos in intact eggs using Laser Speckle Contrast Imaging (LSCI) and Deep Neural Network (DNN)
Simon Mahler,
Anika Arora,
Carol Readhead,
Siyuan Yin,
Surya Narayanan Hari,
Ellie Wang,
Cecilia I Moxley,
Abdullahi A Adeboye
et al.
JY4A4UX3
journalArticle
Jackson Nyman,
Thomas Denize,
Ziad Bakouny,
Chris Labaki,
Breanna M. Titchen,
Kevin Bi,
Surya Narayanan Hari,
Jacob Rosenthal
et al.
ANIMUGTR
journalArticle
Jacob Rosenthal,
Ryan Carelli,
Mohamed Omar,
David Brundage,
Ella Halbert,
Jackson Nyman,
Surya N. Hari,
Eliezer M. Van Allen
et al.
Abstract
Imaging datasets in cancer research are growing exponentially in both quantity and information density. These massive datasets may enable derivation of insights for cancer research and clinical care, but only if researchers are equipped with the tools to leverage advanced computational analysis approaches such as machine learning and artificial intelligence. In this work, we highlight three themes to guide development of such computational tools: scalability, standardization, and ease of use. We then apply these principles to develop PathML, a general-purpose research toolkit for computational pathology. We describe the design of the PathML framework and demonstrate applications in diverse use cases. PathML is publicly available at www.pathml.com.
5F8FLJ28
preprint
Jialong Jiang,
David A. Sivak,
Matt Thomson
The inverse statistical problem of finding direct interactions in complex networks is difficult. In the natural sciences, well-controlled perturbation experiments are widely used to probe the structure of complex networks. However, our understanding of how and why perturbations aid inference remains heuristic, and we lack automated procedures that determine network structure by combining inference and perturbation. Therefore, we propose a general mathematical framework to study inference with iteratively applied perturbations. Using the formulation of information geometry, our framework quantifies the difficulty of inference and the information gain from perturbations through the curvature of the underlying parameter manifold, measured by Fisher information. We apply the framework to the inference of spin network models and find that designed perturbations can reduce the sampling complexity by $10^6$-fold across a variety of network architectures. Physically, our framework reveals that perturbations boost inference by causing a network to explore previously inaccessible states. Optimal perturbations break spin-spin correlations within a network, increasing the information available for inference and thus reducing sampling complexity by orders of magnitude. Our active learning framework could be powerful in the analysis of complex networks as well as in the rational design of experiments.
USG4B25T
computerProgram
A. V. Aditya
A plug-in debugger and visualizer for RL reward functions. Detects reward hacking, tracks training health, and renders a live terminal dashboard.
QKQ2D7C6
computerProgram
Prabhat Prakash
Tutorial to run basic DFT, MD and surface science calculations, in a classroom. Implementing minimal tools, lite exposure to pyscf, ORCA, Jaguar, lammps, Gromacs, Quantum Espresso and VASP codes
YGA4ZWFY
journalArticle
Adri C. T. van Duin,
Siddharth Dasgupta,
Francois Lorant,
William A. Goddard
YW4WQ5KG
computerProgram
syt
Download PDF from Sci-Hub automatically For Zotero7
C4PTEUMI
journalArticle
Zhuoran Qiao,
Matthew Welborn,
Animashree Anandkumar,
Frederick R. Manby,
Thomas F. Miller
We introduce a machine learning method in which energy solutions from the Schrodinger equation are predicted using symmetry adapted atomic orbitals features and a graph neural-network architecture. \textsc{OrbNet} is shown to outperform existing methods in terms of learning efficiency and transferability for the prediction of density functional theory results while employing low-cost features that are obtained from semi-empirical electronic structure calculations. For applications to datasets of drug-like molecules, including QM7b-T, QM9, GDB-13-T, DrugBank, and the conformer benchmark dataset of Folmsbee and Hutchison, \textsc{OrbNet} predicts energies within chemical accuracy of DFT at a computational cost that is thousand-fold or more reduced.
4NAKXPB2
preprint
Zhangde Song,
Jieyu Lu,
Yuanqi Du,
Botao Yu,
Thomas M. Pruyn,
Yue Huang,
Kehan Guo,
Xiuzhe Luo
et al.
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific "superintelligence". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.
56H7SUPA
preprint
Zhangde Song,
Jieyu Lu,
Yuanqi Du,
Botao Yu,
Thomas M. Pruyn,
Yue Huang,
Kehan Guo,
Xiuzhe Luo
et al.
Large language models (LLMs) are increasingly applied to scientific research, yet prevailing science benchmarks probe decontextualized knowledge and overlook the iterative reasoning, hypothesis generation, and observation interpretation that drive scientific discovery. We introduce a scenario-grounded benchmark that evaluates LLMs across biology, chemistry, materials, and physics, where domain experts define research projects of genuine interest and decompose them into modular research scenarios from which vetted questions are sampled. The framework assesses models at two levels: (i) question-level accuracy on scenario-tied items and (ii) project-level performance, where models must propose testable hypotheses, design simulations or experiments, and interpret results. Applying this two-phase scientific discovery evaluation (SDE) framework to state-of-the-art LLMs reveals a consistent performance gap relative to general science benchmarks, diminishing return of scaling up model sizes and reasoning, and systematic weaknesses shared across top-tier models from different providers. Large performance variation in research scenarios leads to changing choices of the best performing model on scientific discovery projects evaluated, suggesting all current LLMs are distant to general scientific "superintelligence". Nevertheless, LLMs already demonstrate promise in a great variety of scientific discovery projects, including cases where constituent scenario scores are low, highlighting the role of guided exploration and serendipity in discovery. This SDE framework offers a reproducible benchmark for discovery-relevant evaluation of LLMs and charts practical paths to advance their development toward scientific discovery.
JTC7P7A5
journalArticle
Yangrui Hu,
Yi Hong Teoh,
William Witczak-Krempa,
Roger G Melko
Abstract
We explore the interplay of quantum computing and machine learning to advance experimental protocols for observing measurement-induced phase transitions (MIPTs) in quantum devices. In particular, we focus on trapped ion monitored circuits and apply the cross entropy benchmark recently introduced by Li
et al
(2023
Phys. Rev. Lett.
130
220404), which can mitigate the post-selection problem. By doing so, we reduce the number of projective measurements—the sample complexity—required per random circuit realization, which is a critical limiting resource in real devices. Since these projective measurement outcomes form a classical probability distribution, they are suitable for learning with a standard machine learning generative model. In this paper, we use a recurrent neural network to learn a representation of the measurement record for a native trapped-ion MIPT, and show that using this generative model can substantially reduce the number of measurements required to accurately estimate the cross entropy. This illustrates the potential of combining quantum computing and machine learning to overcome practical challenges in realizing quantum experiments.
J58LCN6R
webpage
Thinking Machines Lab
On-policy, dense supervision is a useful tool for distillation
PIMKECYV
webpage
lukeprog's profile on LessWrong — A community blog devoted to refining the art of rationality
AJPT6PA4
journalArticle
Gwern
AIs limited to pure computation (Tool AIs) supporting humans, will be less intelligent, efficient, and economically valuable than more autonomous reinforcement-learning AIs (Agent AIs) who act on their own and meta-learn, because all problems are reinforcement-learning problems.
IW8CSLPB
webpage
Jiefang Xiao,
Maolin Gao,
Simon Weber,
Guandao Yang,
Daniel Cremers
Learning mappings between infinite-dimensional function spaces, or operator learning, is essential for many machine learning applications. Although transformer-based operators are popular, they often rely on token-wise attention. These methods treat continuous fields as discrete tokens and usually ignore the global functional structure. We introduce \emph{Functional Attention}, which reinterprets attention as a functional correspondence between adaptive bases. Inspired by geometric functional maps, our method replaces softmax affinities with structured linear operators. This yields a compact, generalizable, resolution-invariant representation that explicitly captures global dependencies. Experiments demonstrate that \emph{Functional Attention} can match state-of-the-art performance in many operator learning tasks, including solving PDEs, 3D segmentation, and regression, while remaining robust to varying discretizations. Project page is available at https://github.com/xjffff/FUNCATTN.
RNKDCMCD
preprint
Terence Parr,
Jeremy Howard
This paper is an attempt to explain all the matrix calculus you need in order to understand the training of deep neural networks. We assume no math knowledge beyond what you learned in calculus 1, and provide links to help you refresh the necessary math where needed. Note that you do not need to understand this material before you start learning to train and use deep learning in practice; rather, this material is for those who are already familiar with the basics of neural networks, and wish to deepen their understanding of the underlying math. Don't worry if you get stuck at some point along the way---just go back and reread the previous section, and try writing down and working through some examples. And if you're still stuck, we're happy to answer your questions in the Theory category at forums.fast.ai. Note: There is a reference section at the end of the paper summarizing all the key matrix calculus rules and terminology discussed here. See related articles at http://explained.ai
ZCB58J48
webpage
A collaborative AI workspace, built on your company context. Build and orchestrate agents right alongside your team's projects, meetings, and connected apps.
ZTMITKGH
webpage
A collaborative AI workspace, built on your company context. Build and orchestrate agents right alongside your team's projects, meetings, and connected apps.
RSK9F7ZW
forumPost
Philip Kiely [@philipkiely]
JYCZY28D
blogPost
Freddie deBoer
my descendants can get a job
MNPPIZFJ
journalArticle
Giacomo Torlai,
Guglielmo Mazzola,
Juan Carrasquilla,
Matthias Troyer,
Roger Melko,
Giuseppe Carleo
The experimental realization of increasingly complex synthetic quantum systems calls for the development of general theoretical methods, to validate and fully exploit quantum resources. Quantum-state tomography (QST) aims at reconstructing the full quantum state from simple measurements, and therefore provides a key tool to obtain reliable analytics. Brute-force approaches to QST, however, demand resources growing exponentially with the number of constituents, making it unfeasible except for small systems. Here we show that machine learning techniques can be efficiently used for QST of highly-entangled states, in both one and two dimensions. Remarkably, the resulting approach allows one to reconstruct traditionally challenging many-body quantities - such as the entanglement entropy - from simple, experimentally accessible measurements. This approach can benefit existing and future generations of devices ranging from quantum computers to ultra-cold atom quantum simulators.
7PJIQEGD
webpage