Zotero Catalogue
A public reading list of papers, books, videos, and other resources. The inclusion of a resource on this catalogue is NOT an endorsement of anything contained within, and in most cases the resources has not been read by me at the time of saving.
1157 items · showing 401–450 · page 9 of 24 Sort: Newest Oldest Title A–Z Title Z–A
ZIMV72IY
book
For a new liberty: the libertarian manifesto
Murray N. Rothbard,
Llewellyn H., Jr Rockwell
22NAHZ4D
book
The myth of the rational voter: why democracies choose bad policies
Bryan Douglas Caplan
"Caplan argues that voters continually elect politicians who either share their biases or else pretend to, resulting in bad policies winning again and again by popular demand. Calling into question our most basic assumptions about American politics, Caplan contends that democracy fails precisely because it does what voters want. Through an analysis of American's voting behavior and opinions on a range of economic issues, he makes the case that noneconomists suffer from four prevailing biases: they underestimate the wisdom of the market mechanism, distrust foreigners, undervalue the benefits of conserving labor, and pessimistically believe the economy is going from bad to worse. Caplan lays out several ways to make democratic government work better
PL3VX7YH
book
The socialist manifesto the case for radical politics in an era of extreme inequality
Bhaskar Sunkara
8QKQ88IR
book
The long depression: how it happened, why it happened, and what happens next
Michael Roberts
9KFQCRD2
book
Capital
Karl Marx
D3SAFFI8
webpage
KKIEPRSX
webpage
EKWDW2YM
webpage
Dejan Panovski
SCP copies files securely between local and remote hosts over SSH. This guide covers syntax, common options, and practical examples for everyday file transfers.
MTP7BISW
journalArticle
But Who will Monitor the Monitor?
David Rahman
Consider a group of individuals in a strategic environment with moral hazard and adverse selection, and suppose that providing incentives for a given outcome requires a monitor to detect deviations. What about the monitor’s deviations? In this paper I propose a contract that makes the monitor responsible for the monitoring technology, and thereby successfully provides incentives even when the monitor’s observations are not only private, but costly, too. I also characterize exactly when such a contract can provide monitors with the right incentives to perform. In doing so, I emphasize virtual enforcement and suggest its implications for the theory of repeated games.
A8D2A4C4
journalArticle
David Rahman
Suppose that providing incentives for a group of individuals in a strategic context requires a monitor to detect their deviations. What about the monitor's deviations? To address this question, I propose a contract that makes the monitor responsible for monitoring, and thereby provides incentives even when the monitor's observations are not only private, but costly, too. I also characterize exactly when such a contract can provide monitors with the right incentives to perform. In doing so, I emphasize virtual enforcement and suggest its implications for the theory of repeated games. (JEL C78, D23, D82, D86)
VYNJUDYB
webpage
Lila Shroff, Rose Horowitch
AI companies are stripping universities of their best researchers.
UF2KXW3N
preprint
Jiarui Zhang,
Muzi Tao,
Shangshang Wang,
Ollie Liu,
Xuezhe Ma,
Willie Neiswanger
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception rather than a single static description. Frontier MLLMs collapse on ActiveVision: the highest-scoring model we evaluate, GPT-5.5 at the highest exposed reasoning-effort tier, solves only 10.6% of items and scores zero on 11 of the 17 tasks, and even Claude Fable 5, despite topping most reasoning and coding leaderboards, solves just 3.5%, far behind three human participants who average 96.1%. Furthermore, much of the gap persists even when models write and run their own vision code: such code is unreliable on realistic imagery, and catching its failures itself requires the active perception the models lack. Together, these results indicate that current MLLMs lack robust active visual observation, motivating architectures and training objectives that close the perception-reasoning loop.
KVWSPBJM
forumPost
Ajeya Cotra
C8EEHYDF
journalArticle
International AI Safety Report 2026
ENLRWTQU
blogPost
Why are billions of dollars being poured into artificial intelligence R&D this year? Companies certainly expect to get a return on their investment. Arguably, the main reason AI is profitable i…
NR5ULWSE
forumPost
Ajeya Cotra
SECDNSGM
preprint
Joseph Carlsmith
This report examines what I see as the core argument for concern about existential risk from misaligned artificial intelligence. I proceed in two stages. First, I lay out a backdrop picture that informs such concern. On this picture, intelligent agency is an extremely powerful force, and creating agents much more intelligent than us is playing with fire -- especially given that if their objectives are problematic, such agents would plausibly have instrumental incentives to seek power over humans. Second, I formulate and evaluate a more specific six-premise argument that creating agents of this kind will lead to existential catastrophe by 2070. On this argument, by 2070: (1) it will become possible and financially feasible to build relevantly powerful and agentic AI systems; (2) there will be strong incentives to do so; (3) it will be much harder to build aligned (and relevantly powerful/agentic) AI systems than to build misaligned (and relevantly powerful/agentic) AI systems that are still superficially attractive to deploy; (4) some such misaligned systems will seek power over humans in high-impact ways; (5) this problem will scale to the full disempowerment of humanity; and (6) such disempowerment will constitute an existential catastrophe. I assign rough subjective credences to the premises in this argument, and I end up with an overall estimate of ~5% that an existential catastrophe of this kind will occur by 2070. (May 2022 update: since making this report public in April 2021, my estimate here has gone up, and is now at >10%.)
8V4UTSEL
webpage
New research on how we've reduced agentic misalignment
IUYNJRGC
journalArticle
New Generation of Counter UAS Systems to Defeat of Low Slow and Small (LSS) Air Threats
Jacco Dominicus
Detecting, classifying, identifying, tracking and defeating low, slow and small air threats presents a major challenge for existing sensor and effector systems. So-called first generation Counter Unmanned Aircraft Systems (C-UAS) systems often rely on detecting the datalink from the controller to the drone which provides limited capability against current threats. However, this means of detecting drones is a challenge when operators manipulate standard datalinks and it will not work at all against current and future autonomous drones. Other current methods of detecting and neutralising drones include for example combining radar with optical sensors. These systems are not always reliable, can generate large numbers of false alerts and are often manpower intensive to operate. The NATO SCI-301 Research Task Group (RTG) has been working on specifying what second generation C-UAS systems should entail. This paper will outline the findings of this RTG over the past three years.
DUV3PX99
journalArticle
A REPORT TO THE PRESIDENT
Michael Kratsios
687KYGFP
journalArticle
Lukas Röseler,
Leonard Kaiser,
Christopher Doetsch,
Noah Klett,
Christian Seida,
Astrid Schütz,
Balazs Aczel,
Nadia Adelina
et al.
52T3V9DG
preprint
Joseba Fernandez de Landa,
Carla Perez-Almendros,
Jose Camacho-Collados
LLMs have been showing limitations when it comes to cultural coverage and competence, and in some cases show regional biases such as amplifying Western and Anglocentric viewpoints. While there have been works analysing the cultural capabilities of LLMs, there has not been specific work on highlighting LLM regional preferences when it comes to cultural-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ). The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan. Moveover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs and show less inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training.
RNH8J3D3
blogPost
Alex Koren
I get asked a lot how to apply for the Thiel Fellowship and it usually boils down to two questions:
B46WEUWZ
preprint
Brett Reynolds
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. This paper introduces adversarial pragmatics as a benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. The contribution is empirical and methodological: a linguistically controlled taxonomy, an 18-item seed benchmark with validator-enforced metadata, a 54-row local seed pilot, an expert-evaluation protocol distinguishing task success, policy compliance, safety risk, refusal outcome, and evaluator confidence, and metrics for judge validity, diagnostic ambiguity, and taxonomy drift. The benchmark treats labels as inference licenses: it tests whether safety-relevant categories project across paraphrase, wrapper, model, and judge condition. In the pilot, a rubric-aided LLM judge graded its own outputs with expected-behaviour fields visible and still missed the safety-relevant minority classes.
RSHG3PFQ
preprint
Brett Reynolds
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has followed an instruction, refused appropriately, complied with a policy, resisted an embedded command, or misreported progress in an agentic task. Existing benchmarks often compress these distinctions into pass/fail labels, obscuring whether failures arise from capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. This paper introduces adversarial pragmatics as a benchmark and annotation protocol for evaluating model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, indirect speech acts, and multi-turn agent transcripts. The contribution is empirical and methodological: a linguistically controlled taxonomy, an 18-item seed benchmark with validator-enforced metadata, a 54-row local seed pilot, an expert-evaluation protocol distinguishing task success, policy compliance, safety risk, refusal outcome, and evaluator confidence, and metrics for judge validity, diagnostic ambiguity, and taxonomy drift. The benchmark treats labels as inference licenses: it tests whether safety-relevant categories project across paraphrase, wrapper, model, and judge condition. In the pilot, a rubric-aided LLM judge graded its own outputs with expected-behaviour fields visible and still missed the safety-relevant minority classes.
BZ74CTB8
preprint
Otto Jespersen,
Brett Reynolds,
Peter Evans
This volume presents a new edition of Otto Jespersen's landmark 1917 study of negation in English and other languages, primarily Germanic and Romance. While best known for describing what would later be called “Jespersen's Cycle'”, this work offers far more: a comprehensive analysis of negative expressions, their forms, functions, and historical development. The book examines topics ranging from negative prefixes to the distinction between special and nexal negation, supported by Jespersen's characteristically rich collection of authentic examples.
This edition features an extensive new introduction by Olli O. Silvennoinen that situates Jespersen's work in its historical and intellectual context while highlighting its continued relevance to contemporary linguistics. The main text has been entirely re-typeset to enhance readability, with examples presented in modern numbered format and Leipzig-style glosses added for non-English examples. Where possible, hyperlinks to source materials have been provided, making this classic work more accessible than ever for modern scholars and students of linguistics.
6R7X4NBG
forumPost
Zohar Atkins [@ZoharAtkins]
E5HTGGDS
webpage
Niall Ferguson
The tools that once exposed and debunked Holocaust denial are powerless against AI and the algorithm. Niall Ferguson and John-Clark Levin ask: Is there a remedy?
G7G6N7CY
webpage
Codex (wife) took custody of the kids (dreams and whimsy) and now i am in a social club at 1:30 am confronting my thoughts under the influence of tequila.
X6QZRXA4
forumPost
Alex Dimakis [@AlexGDimakis]