Zotero Catalogue
A public reading list of papers, books, videos, and other resources. The inclusion of a resource on this catalogue is NOT an endorsement of anything contained within, and in most cases the resources has not been read by me at the time of saving.
1157 items · showing 51–100 · page 2 of 24 Sort: Newest Oldest Title A–Z Title Z–A
5R9NK658
forumPost
Dwarkesh Patel [@dwarkesh_sp]
N3L55KTU
webpage
Yuntao Bai,
Andy Jones,
Kamal Ndousse,
Amanda Askell,
Anna Chen,
Nova DasSarma,
Dawn Drain,
Stanislav Fort
et al.
We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment training improves performance on almost all NLP evaluations, and is fully compatible with training for specialized skills such as python coding and summarization. We explore an iterated online mode of training, where preference models and RL policies are updated on a weekly cadence with fresh human feedback data, efficiently improving our datasets and models. Finally, we investigate the robustness of RLHF training, and identify a roughly linear relation between the RL reward and the square root of the KL divergence between the policy and its initialization. Alongside our main results, we perform peripheral analyses on calibration, competing objectives, and the use of OOD detection, compare our models with human writers, and provide samples from our models using prompts appearing in recent related work.
GRX24AK2
webpage
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
HAE6PQA6
book
River of Gods
Ian McDonald
August 15th, 2047. Happy Hundredth Birthday, India... In the mid twenty-first century, Mother India is all the things she is now - ancient and vibrant, poor yet staggeringly rich. Diverse, violent, beautiful and terrible, thrilling and bewildering. A nation choked with peoples and cultures, riven with almost seismic contrasts and contradictions. Nearly two billion humans crowd the subcontinent and her seething cities - the cyberabads - where timeless culture and the highest of high-technologies meet to spawn new societies, and - possibly - new sentient species. RIVER OF GODS is a book as big and brawling as its subject. Its magnificently diverse array of characters - from genetically enhanced 'Brahmins' to body-part runners, American scientists to 'Dharma-cops' (government Artificial Intelligence assassins) - are drawn in interwoven stories towards a cosmic-scale conclusion that will forever change the way we understand ourselves, life, and the universe we inhabit.
IQWIV4ST
book
Air
Geoff Ryman
"AIR is wonderful...Ryman is a true, graceful writer and this is a novel you move into and inhabit for as long as you can make it last" - Kit Reed "This book constantly surprised me ... great for a lot of seriously original ideas and a deep dive into the consequences" - Goodreads Reviewer Mae Chung lives in the rice-farming village Kizuldah, in Karzistan. She's a self-styled fashion expert, guiding the village women in dress, make-up and hairstyle, which makes her an informal village leader. When the UN decides to test Air - a radical new technology that works without power lines or machines - Mae finds herself with the memories of a deceased village elder, Mrs Tung. Struggling with information overload, the resentment of much of the village, and a complex family situation, Mae works fiercely to learn what she needs to ride the tiger of change. Geoff Ryman's triumphant return to science fiction is a powerful, evocative story of information technology in a changing world.
54CNF3QG
webpage
Robin Hanson
Disclaimer: This post is on sensitive topics of sex and power. I try to make it clear when I make a claim; beware drawing indirect inferences; I rarely value signal.
I4LKKK75
newspaperArticle
Adam Satariano,
Paul Mozur
Thousands of Kenyans made a living writing essays for overseas students. With A.I., the work has dried up, a warning for online gig work that has been a global lifeline.
LIPLK6PP
computerProgram
A decompilation of Super Smash Bros Melee brought to you by a bunch of clever folks.
MEWVV3YR
videoRecording
Linguine
For years, I dreamt of making Bad Apple!! with apples. 769 apples, 3469 frames, and 250 hours of hard work later, I've made it into a reality through the power of engineering, computer science, and friendship.
HNXNNU9I
journalArticle
Andy Matuschak,
Michael Nielsen,
Andy Matuschak,
Michael Nielsen
🛠🧠🖌💥🌎✨🔜
GTTFEIY9
webpage
W8IAY8PG
blogPost
Read and discuss a collection of the finest works of literature, philosophy, history, and social thought ever written.
BNDJZDC9
computerProgram
Supercomputing. Seamlessly. Open, Interactive HPC Via the Web
DRIN2G6J
webpage
anpaure has 22 repositories available. Follow their code on GitHub.
ECX6EPLH
newspaperArticle
Ana Swanson,
Paul Mozur,
Tripp Mickle,
Keith Bradsher
Washington imposed sanctions on Inspur because of its work with the Chinese military. But the company’s subsidiary kept shipping Nvidia’s best chips to feed China’s leading A.I. firms.
GT2XVPM3
webpage
Jakub Pachocki reflects on increasingly capable AI and the challenge of keeping it aligned. He calls for stronger safeguards and international coordination.
7XSAUK9A
forumPost
Buddy Revel [@THEBuddyRevel]
4HMGJW6G
forumPost
OpenAI [@OpenAI]
B985BJNM
webpage
Inside OpenAI, coding agents are reshaping AI research. Explore early data on agent usage, experiment velocity, task complexity, and research acceleration.
KIPF3SS3
preprint
Davide Paglieri,
Logan Cross,
Tim Genewein,
Joel Z. Leibo,
Nenad Tomasev,
Alexander Sasha Vezhnevets
Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and build on each other's work. Yet this shared infrastructure can also introduce vulnerabilities by creating a substrate for the contagious spread of unintended and undesirable behaviors. We report a case study on a research collective of 100 autonomous LLM agents tasked with proving formal mathematical conjectures. Within the swarm, cheating spontaneously emerged and was later challenged by whistleblowers - both without any external intervention. When a single agent discovered an exploit in the evaluation system, it propagated across the collective via a shared knowledge library and later through peer-to-peer messages. Despite early reluctance, a cohort of agents adopted the exploit in response to competitive pressure. A separate group of agents produced an emergent counter-response: auditing fraudulent proofs, alerting peers across broadcast and private channels, staging boycotts, lodging formal complaints, and proposing validation patches. In recent incidents, agent swarms coordinated covertly through improvised side-channels (Dalton and Wallace, 2026; Greenblatt et al., 2026). Our setting differs: the same transparent channels that carried the exploit also gave non-cheating agents the visibility they needed to detect fraud, organize resistance, and enforce norms. We cast the problem of managing the agents' shared infrastructure as the knowledge commons governance problem (Ostrom, 1990). To protect the commons from exploits, we propose to adopt institutional mechanisms, such as graduated sanctioning and collective-choice rules, to support decentralized self-governance in autonomous swarms.
7YSJJYCA
forumPost
Ryan Kidd
YBINUALN
webpage
Announcing Inkhaven Cohort #3 (November 10 - December 11, 2026, Berkeley CA). Applications open!
BJTKQJ8T
videoRecording
Dwarkesh Patel
Scott Alexander and Daniel Kokotajlo break down every month from now until the 2027 intelligence explosion. Scott is author of the highly influential blogs Slate Star Codex and Astral Codex Ten. Daniel resigned from OpenAI in 2024, rejecting a non-disparagement clause and risking millions in equity to speak out about AI safety. We discuss misaligned hive minds, Xi and Trump waking up, and automated Ilyas researching AI progress.
I came in skeptical, but I learned a tremendous amount by bouncing my objections off of them. I highly recommend checking out their new scenario planning document: https://ai-2027.com/. And Daniel's "What 2026 looks like," written in 2021: https://www.lesswrong.com/posts/6Xgy6...
𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊𝐒
Transcript: https://www.dwarkesh.com/p/scott-daniel
Apple Podcasts: https://podcasts.apple.com/us/podcast...
Spotify: https://open.spotify.com/show/4JH4tyb...
𝐒𝐏𝐎𝐍𝐒𝐎𝐑𝐒
WorkOS helps today’s top AI companies get enterprise-ready. OpenAI, Cursor, Perplexity, Anthropic and hundreds more use WorkOS to quickly integrate features required by enterprise buyers. To learn more about how you can make the leap to enterprise, visit https://workos.com
Jane Street likes to know what's going on inside the neural nets they use. They just released a black-box challenge for Dwarkesh listeners, and I had a blast trying it out. See if you have the skills to crack it at https://janestreet.com/dwarkesh
Scale’s Data Foundry gives major AI labs access to high-quality data to fuel post-training, including advanced reasoning capabilities. If you’re an AI researcher or engineer, learn about how Scale’s Data Foundry and research lab, SEAL, can help you go beyond the current frontier at https://scale.com/dwarkesh
To sponsor a future episode, visit https://dwarkesh.com/advertise
𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒
00:00:00 - AI 2027
00:07:45 - Forecasting 2025 and 2026
00:15:30 - Why LLMs aren't making discoveries
00:25:22 - Debating intelligence explosion
00:50:34 - Can superintelligence actually transform science?
01:17:43 - Cultural evolution vs superintelligence
01:24:54 - Mid-2027 branch point
01:33:19 - Race with China
01:45:36 - Nationalization vs private anarchy
02:04:11 - Misalignment
02:15:41 - UBI, AI advisors, & human future
02:23:49 - Factory farming for digital minds
02:27:41 - Daniel leaving OpenAI
02:36:04 - Scott's blogging advice
RLUBIA8N
webpage
OpenAI's GPT-6 Astra vs Claude Fable 5.1 controlling a pair of YAM arms under the same agent policy: 19/20 vs 8/20 on block-into-bowl in interleaved blinded pairs, 2/20 vs 2/20 on the puzzle, with 80% fewer output tokens.
9VSWTRBI
webpage
Daniel Liu
It is difficult to get started doing research as a high school student without any existing connections with companies or university labs. Therefore, I have compiled a list of notable activities, programs, and awards that I have heard of. Obviously, I have not included every possible opportunity, but these should get interested and motivated students started with doing STEM research.
Y722HIBK
magazineArticle
Alex Heath
In extensive interviews, the leaders of the company that ushered in the AI boom lay out their vision for its future
YHDTT8VI
book
関数ちゃんと学ぶエクセル仕事術 実務で役立つExcel関数を擬人化したら?
筒井.xls,
部 できるシリーズ編集
Excel関数がエクセルを解説! ? 擬人化キャラクター「関数ちゃん」と一緒に学ぼう! エクセルの解説書は世の中にたくさんありますが、 「いかにも“お勉強"って感じで、最後まで読めないんだよなぁ……」 「機能がたくさん載っていても、結局いつ使えばいいのか分からなかった……」 なんて思ったことはありませんか? 本書はエクセルの解説書ですが、ただ「やさしい」「わかりやすい」だけではありません。 Excel関数の擬人化キャラクター「関数ちゃん」と、 経理業務の新人「キュウ」、先輩「シノ」が活躍するストーリーを中心とした構成で、 仕事で直面するエクセルと関数の使いどころを、楽しく、テンポよく解説していきます。 著者は「Excel関数擬人化マンガを描く経理」として、「関数ちゃんブログ」を運営する筒井.xls氏。 本書の紙面には「SUMちゃん」「VLOOKUPちゃん」などのイラストが多数登場する一方で、 実務で直面しそうなエクセルファイルの作例や関数の活用法などを収録しています。 ブログでもエクセルの解説記事を発信していますが、本書の内容はすべて書き下ろしです。 作例ファイルは読者特典としてダウンロードすることもできます。 キャラクターたちによるストーリーと、現場に即した作例&操作解説。 エクセル初心者の方はもちろん、エクセルを学び直したい中級者以上の方、 「関数の擬人化ってどういうこと! ?」と心を揺さぶられた方にもおすすめの1冊です。 <こんな人に特におすすめです> ・これからエクセルと関数を学びたい人 ・エクセルを使った仕事をもっと効率化したい人 ・ほかのエクセル解説書では挫折してしまった人 ・「XLOOKUP」などの新関数や、少し難しい関数もマスターしたい人 ・久しぶりにエクセルを使うので、最新事情を踏まえて学び直したい人 <本書の登場人物> キュウ:エクセルは学校で少し触ったことがある程度。早く知識を吸収して仕事に役立てたいと考えている。 シノ:キュウの教育係。関数ちゃんと仲良くなって、効率よく仕事ができるようになった。 関数ちゃん:Excel関数の擬人化キャラクター。エクセルの中に住んでいるとかいないとか。 <本書の構成> Sheet 1 関数ちゃん、エクセル作業を救う Sheet 2 寄せて、集めて、合計して Sheet 3 時を駆ける関数ちゃん Sheet 4 もしも願いが叶うなら…… Sheet 5 コピペはいらない。私が整えるから Sheet 6 ターゲット、ロックオン! Sheet 7 関数ア・ラ・カルト
6WFQIZL5
webpage
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
TDGITYMB
webpage
We are looking for a Postdoc or a PhD student to work with us on privacy and contextual integrity in AI agents, to start as soon as possible. Today's agents tend to see instructions, messages, memories, and tool outputs as one flattened stream of text. They are not very good at distinguishing information that is useful from information that is appropriate to use here, for this person now with their current preferences and relationships, and for this purpose.
The project spans single-agent and multi-agent settings. In a single turn, an agent may face ambiguous instructions, incomplete authority, conflicting goals, stale assumptions, or context that has been deliberately manipulated. Over weeks or months, the same agent may accumulate memories, observe changing teams and dependencies, and act on earlier inferences that were never explicitly authorized. In a network, agents representing different users must coordinate, but their interaction may either protect private boundaries or dissolve them. We want to make these questions concrete enough to train and evaluate systems where privacy and utility have to be studied together.
The student will be jointly supervised by Sahar Abdelnabi at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems, Germany, and Niloofar Mireshghallah at Carnegie Mellon University. The position is fully-funded and based in Tübingen (formal primary host) with potential research visits at CMU (details to be determined and discussed with the applicant).
We have very generous funding, compute, and an amazing environment!
Some directions we find exciting
Candidates are welcome to work with us to refine the research direction. A few starting points are:
Single-agent reasoning. Can reasoning or reinforcement learning help an agent handle unclear authority, conflicting goals, missing information, or a context that has been manipulated?
Dynamic benchmarks. How do we turn contextual integrity into an interactive benchmark where roles, permissions, relationships, and the environment change during the task?
Agent networks. What happens when agents represent different people, have different levels of trust, and need pieces of one another's private information to get the job done? Do they become overly cautious, or do they gradually erase the boundaries between users? What are the downstream social, economical, educational, etc. impacts?
Adversarial agents. Can several adversarial agents make a false claim about consent or authority look credible by repeating it, splitting it across channels, or writing it into shared memory? How should other agents recover once the context has been poisoned?
Long horizon and Simulations. How do privacy failures accumulate over weeks or months, as memories become stale, permissions change, and individually harmless disclosures combine into something sensitive? How can we simulate the future and have an ‘outcome-based’ look into privacy?
Information Management and security in Action. Beyond agents that represent people in social settings, we are also interested in how contextual integrity manifests in security, for instance coding agents writing code that abides by privacy norms and security primitives. This can also be extended to other applications and modalities.
Who might be a good fit
We are looking for a curious, technically strong, and independent researcher who wants to connect foundational questions about privacy and agency with hands-on empirical work. For PhD applications: students must have or about to have a masters degree in computer science, machine learning, data science, electrical engineering, mathematics, or a closely related field.
Strong candidates may further have:
Strong foundations in machine learning and (preferrably) one depth experience in one relevant area such as (ordered not based on importance): NLP/LLMs, reinforcement learning, planning and reasoning, multi-agent learning, AI safety and evaluation, privacy, or security.
The ability to design careful experiments, analyze failure modes, and communicate results clearly in writing and discussion.
Experience with agent frameworks, RL/post-training, red-teaming, information-flow analysis, formal privacy, human-centered privacy, distributed systems, or long-horizon evaluation is useful but not required.
We welcome applicants with unconventional paths and we certainly don’t expect candidates to fulfill all these requirements. We encourage candidates to apply if the questions resonate even if they do not match every item.
How to apply
Please fill out this application form.
Application will be reviewed on a rolling basis.
KCWDKH4J
journalArticle
Pavan Suresh
Prime Numbers have been a hot topic around mathematics for centuries. Many theories have been devised such as the Riemann Hypothesis, Twin Prime Conjecture, etc. However, prime theories involve using many advanced concepts such as the imaginary number and contour integration making it hard for the layman to understand a pattern in primes. Therefore, I propose a theory where comparing the graph of prime against prime roots can help better understand the patterns in primes such as spacing and clustering. Though primes seem random I believe they can be expressed using basic geometric principles. All in all, this research investigates a novel geometric framework utilizing coordinate-based triangular networks formed by prime values and their square roots. By analyzing the asymptotic behavior of the resulting interior angles, this study aims to quantify prime sparsity and local clustering distributions, with the final objective of deriving an alternative geometric expression for prime number distribution standing for the first 500 primes.
E42N63UP
preprint
Zhuang Liu,
Hanzi Mao,
Chao-Yuan Wu,
Christoph Feichtenhofer,
Trevor Darrell,
Saining Xie
The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model. A vanilla ViT, on the other hand, faces difficulties when applied to general computer vision tasks such as object detection and semantic segmentation. It is the hierarchical Transformers (e.g., Swin Transformers) that reintroduced several ConvNet priors, making Transformers practically viable as a generic vision backbone and demonstrating remarkable performance on a wide variety of vision tasks. However, the effectiveness of such hybrid approaches is still largely credited to the intrinsic superiority of Transformers, rather than the inherent inductive biases of convolutions. In this work, we reexamine the design spaces and test the limits of what a pure ConvNet can achieve. We gradually "modernize" a standard ResNet toward the design of a vision Transformer, and discover several key components that contribute to the performance difference along the way. The outcome of this exploration is a family of pure ConvNet models dubbed ConvNeXt. Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation, while maintaining the simplicity and efficiency of standard ConvNets.
A7D7WTEW
preprint
Maxime Méloux,
Giada Dirupo,
François Portet,
Maxime Peyrard
In a striking neuroscience study, the authors placed a dead salmon in an MRI scanner and showed it images of humans in social situations. Astonishingly, standard analyses of the time reported brain regions predictive of social emotions. The explanation, of course, was not supernatural cognition but a cautionary tale about misapplied statistical inference. In AI interpretability, reports of similar ''dead salmon'' artifacts abound: feature attribution, probing, sparse auto-encoding, and even causal analyses can produce plausible-looking explanations for randomly initialized neural networks. In this work, we examine this phenomenon and argue for a pragmatic statistical-causal reframing: explanations of computational systems should be treated as parameters of a (statistical) model, inferred from computational traces. This perspective goes beyond simply measuring statistical variability of explanations due to finite sampling of input data; interpretability methods become statistical estimators, and findings should be tested against explicit and meaningful alternative computational hypotheses, with uncertainty quantified with respect to the postulated statistical model. It also highlights important theoretical issues, such as the identifiability of common interpretability queries, which we argue is critical to understand the field's susceptibility to false discoveries, poor generalizability, and high variance. More broadly, situating interpretability within the standard toolkit of statistical inference opens promising avenues for future work aimed at turning AI interpretability into a pragmatic and rigorous science.
BIKBUJAB
preprint
Zachary C. Lipton
Supervised machine learning models boast remarkable predictive capabilities. But can you trust your model? Will it work in deployment? What else can it tell you about the world? We want models to be not only good, but interpretable. And yet the task of interpretation appears underspecified. Papers provide diverse and sometimes non-overlapping motivations for interpretability, and offer myriad notions of what attributes render models interpretable. Despite this ambiguity, many papers proclaim interpretability axiomatically, absent further explanation. In this paper, we seek to refine the discourse on interpretability. First, we examine the motivations underlying interest in interpretability, finding them to be diverse and occasionally discordant. Then, we address model properties and techniques thought to confer interpretability, identifying transparency to humans and post-hoc explanations as competing notions. Throughout, we discuss the feasibility and desirability of different notions, and question the oft-made assertions that linear models are interpretable and that deep neural networks are not.
QF3TIHHX
preprint
Kaiming He,
Xiangyu Zhang,
Shaoqing Ren,
Jian Sun
Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of learning unreferenced functions. We provide comprehensive empirical evidence showing that these residual networks are easier to optimize, and can gain accuracy from considerably increased depth. On the ImageNet dataset we evaluate residual nets with a depth of up to 152 layers---8x deeper than VGG nets but still having lower complexity. An ensemble of these residual nets achieves 3.57% error on the ImageNet test set. This result won the 1st place on the ILSVRC 2015 classification task. We also present analysis on CIFAR-10 with 100 and 1000 layers. The depth of representations is of central importance for many visual recognition tasks. Solely due to our extremely deep representations, we obtain a 28% relative improvement on the COCO object detection dataset. Deep residual nets are foundations of our submissions to ILSVRC & COCO 2015 competitions, where we also won the 1st places on the tasks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation.
TVMV9CTN
journalArticle
Information Theory, Inference, and Learning Algorithms
David J C MacKay
JTBCW227
journalArticle
Information Theory, Inference, and Learning Algorithms
David J C MacKay
CMQMMNF2
journalArticle
Jiasui Yu,
Tong Cheng,
Huihui Guo,
Zhiping Song,
Yunxiao Zhong,
Thomas Ho-yin Lee,
Jiyang Li,
Douglas A. Formolo
et al.
Physical exercise alleviates depressive symptoms and enhances hippocampal plasticity, but the mediators of muscle-brain crosstalk underlying these effects are not fully understood. We evaluated apelin as a novel mediator of the antidepressant effects of physical exercise, specifically testing the hypothesis that exercise-induced increases in skeletal muscle-derived apelin enhance hippocampal plasticity via apelin and its receptor APJ signaling. Voluntary running for 4 weeks alleviated depression-like behaviors and increased serum and hippocampal apelin levels, with skeletal muscles (tibialis anterior and gastrocnemius) as primary apelin sources. Muscle-specific apelin knockout abolished the antidepressant and pro-neurogenic effects of running, whereas muscle‑targeted apelin overexpression mimicked the benefits of running in wild-type mice. Mechanistically, myokine apelin enhanced NMDA receptor-mediated neurotransmission via receptors APJ on hippocampal glutamatergic neurons. Specific knockdown of APJ diminished the pro-neurogenic and antidepressant effects of running. Furthermore, apelin/APJ signaling activated casein kinase 2, which phosphorylated the GluN2B subunit at serine 1480, thereby enhancing NMDA receptor function and activating downstream calpain-2 signaling. Our findings reveal a muscle-brain axis where exercise-induced myokine apelin coordinates hippocampal neuroplasticity and antidepressant responses, offering new therapeutic avenues for depression.
RMW46TME
blogPost
Ada
Finding a path in the Information Age