Zotero Catalogue
A public reading list of papers, books, videos, and other resources. The inclusion of a resource on this catalogue is NOT an endorsement of anything contained within, and in most cases the resources has not been read by me at the time of saving.
1157 items · showing 751–800 · page 16 of 24 Sort: Newest Oldest Title A–Z Title Z–A
ZRQC5W9K
forumPost
DazzlingPin3965
784VK76U
webpage
Dr Werner Vogels- https://www.allthingsdistributed.com
The subtle inventiveness that reduced cold start setup from seconds to 200μs.
5QT54B2H
journalArticle
Kent C. Berridge,
Terry E. Robinson
Rewards are both ‘liked’ and ‘wanted’, and those two words seem almost interchangeable. However, the brain circuitry that mediates the psychological process of ‘wanting’ a particular reward is dissociable from circuitry that mediates the degree to which it is ‘liked’. Incentive salience or ‘wanting’, a form of motivation, is generated by large and robust neural systems that include mesolimbic dopamine. By comparison, ‘liking’, or the actual pleasurable impact of reward consumption, is mediated by smaller and fragile neural systems, and is not dependent on dopamine. The incentive-sensitization theory posits the essence of drug addiction to be excessive amplification specifically of psychological ‘wanting’, especially triggered by cues, without necessarily an amplification of ‘liking’. This is due to long-lasting changes in dopamine-related motivation systems of susceptible individuals, called neural sensitization. A quarter-century after its proposal, evidence has continued to grow in support the incentive-sensitization theory. Further, its scope is now expanding to include diverse behavioral addictions and other psychopathologies.
6P6NJEQD
webpage
Viola Zhou
Chinese-born tech workers have fueled Silicon Valley for decades. In the AI era, they're superstars.
5D57K3UE
bookSection
Aslak Tveito,
Are Magnus Bruaset,
Olav Lysne,
J. F. Kaiser
W4PWEH5P
forumPost
Neal Parikh [@npparikh]
JHQ2RHTI
preprint
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
Fan Long
KKFWFEG5
preprint
Tara Saba,
Anne Ouyang,
Xujie Si,
Fan Long
High-performance GPU kernels are critical to modern machine learning systems, yet developing efficient implementations remains a challenging, expert-driven process due to the tight coupling between algorithmic structure, memory hierarchy usage, and hardware-specific optimizations. Recent work has explored using large language models (LLMs) to generate GPU kernels automatically, but generated implementations often struggle to maintain correctness and achieve competitive performance across iterative refinements. We present CuTeGen, an agentic framework for automated generation and optimization of GPU kernels that treats kernel development as a structured generate--test--refine workflow. Unlike approaches that rely on one-shot generation or large-scale search over candidate implementations, CuTeGen focuses on progressive refinement of a single evolving kernel through execution-based validation, structured debugging, and staged optimization. A key design choice is to generate kernels using the CuTe abstraction layer, which exposes performance-critical structures such as tiling and data movement while providing a more stable representation for iterative modification. To guide performance improvement, CuTeGen incorporates workload-aware optimization prompts and delayed integration of profiling feedback. Experimental results on matrix multiplication and activation workloads demonstrate that the framework produces functionally correct kernels and achieves competitive performance relative to optimized library implementations.
NZM98V5D
conferencePaper
Fan Long,
Martin Rinard
HSYCIZGZ
preprint
Yuhui Li,
Fangyun Wei,
Chao Zhang,
Hongyang Zhang
The sequential nature of modern LLMs makes them expensive and slow, and speculative sampling has proven to be an effective solution to this problem. Methods like EAGLE perform autoregression at the feature level, reusing top-layer features from the target model to achieve better results than vanilla speculative sampling. A growing trend in the LLM community is scaling up training data to improve model intelligence without increasing inference costs. However, we observe that scaling up data provides limited improvements for EAGLE. We identify that this limitation arises from EAGLE's feature prediction constraints. In this paper, we introduce EAGLE-3, which abandons feature prediction in favor of direct token prediction and replaces reliance on top-layer features with multi-layer feature fusion via a technique named training-time test. These improvements significantly enhance performance and enable the draft model to fully benefit from scaling up training data. Our experiments include both chat models and reasoning models, evaluated on five tasks. The results show that EAGLE-3 achieves a speedup ratio up to 6.5x, with about 1.4x improvement over EAGLE-2. In the SGLang framework, EAGLE-3 achieves a 1.38x throughput improvement at a batch size of 64. The code is available at https://github.com/SafeAILab/EAGLE.
JVC7QYW3
preprint
Hongyang Zhang
NYXB3PTG
blogPost
Matt Giaro
This app is the backbone of my 6-figure writing business
2CEM2ZK5
journalArticle
From Worm to Human: Scaling Brain Emulation
Isaak Freeman
Machine learning models are rapidly approaching or surpassing human performance on many metrics. In comparison, neuroscience is progressing at a slow pace. To reach the research velocity possible in software paradigms, we highlight a potential path to high-quality emulations of the brains of key model organisms such as the roundworm Caenorhabditis elegans, the zebrafish Danio rerio and the mouse Mus musculus.
AGRC2GVW
blogPost
John Friedman
Scaling data processing, starting with the SEC corpus.
3CAUY7QL
blogPost
John Friedman
Addressing a recent, popular misconception
LF43M5BV
blogPost
John Friedman
I had fun.
XGLETZKS
webpage
Ted Leung noted the discussion that Werner and I have been having, and observed that we should consider Rob Pike’s (in)famous polemic, “Systems Software Research is Irrelevant.” I should say that I broadly agree with most of Pike’s conclusions – and academic systems software research has seemed increasingly irrelevant in the last five years. That said, I think that what Pike characterizes as “systems research” is far too skewed to the interface to the system – which (tautologically) is but the periphery of the larger system. In my opinion, “systems research” should focus not on the interface of the system, but rather its guts: those hidden Rube Goldberg-esque innards that are rife with old assumptions and unintended consequences. Pike would perhaps dismiss the study of these innards as “phenomenology”, but I would counter that understanding phenomena is a prerequisite to understanding larger systemic truths. Of course, the problem to date has been that much systems research has not been able to completely understand phenomena – the research has often consisted merely of characterizing it.
MNH47WSR
webpage
Note: Thie was co-authored with Steve Tuck, and originally appeared on the Oxide blog.
We don’t want to bury the lede: we have raised a $100M Series B, led by a new strategic partner in USIT with participation from all existing Oxide investors. To put that number in perspective: over the nearly six year lifetime of the company, we have raised $89M; our $100M Series B more than doubles our total capital raised to date — and positions us to make Oxide the generational company that we have always aspired it to be.
H2RHMNUZ
webpage
Last Tuesday, several months of preparation came to fruition in the inaugural Systems We Love. You never know what’s going to happen the first time you get a new kind of conference together (especially one as broad as this one!) but it was, in a word, amazing. The content was absolutely outstanding, with attendee after attendee praising the uniformly high quality. (For guided tours, check out both Ozan Onay’s excellent exegesis and David Cassel’s thorough New Stack story – and don’t miss Sarah Huffman’s incredible illustrations!) It was such a great conference that many were asking about when we would do it again – and there is already interest in replicating it elsewhere. As an engineer, this makes me slightly nervous as I believe that success often teaches you nothing: luck becomes difficult to differentiate from design. But at the risk of taunting the conference gods with the arrogance of a puny mortal, here’s some stuff I do think we did right:
RQ588BYF
webpage
USENIX made the decision this week to discontinue its flagship Annual Technical Conference. When USENIX was started in 1975 — before the Internet, really — conferences were the fastest vector for practitioners to formally share their ideas, and USENIX ATC flourished. Speaking for myself, I came up lionizing ATC: I was an undergraduate in the early 1990s, and programs like the USENIX Summer 1994 conference felt like Renaissance-era Florence for systems practitioners.
XUI6NAKY
journalArticle
Systems Software Research is Irrelevant
Rob Pike,
Bell Labs
V57EMEKD
webpage
Derek Sivers official site. Thoughts on philosophy, culture, self-improvement. Author of Useful Not True, How to Live, Hell Yeah or No, Anything You Want.
H45DILSI
webpage
Derek Sivers official site. Thoughts on philosophy, culture, self-improvement. Author of Useful Not True, How to Live, Hell Yeah or No, Anything You Want.
CXPT4MHB
webpage
Derek Sivers official site. Thoughts on philosophy, culture, self-improvement. Author of Useful Not True, How to Live, Hell Yeah or No, Anything You Want.
8HIJUU8Z
magazineArticle
James Somers
Coding together at the same computer, Jeff Dean and Sanjay Ghemawat changed the course of the company—and the Internet.
JPPEJZ3L
webpage
Christopher Olah
Instead of asking “Is university good?”, ask “Do I have something more compelling to do?”. Instead of “Should I do a PhD?”, ask “Where can I find the best environment to grow as a researcher?”.
TC3I77TI
webpage
Many founders choose the factory, and they choose it for understandable reasons: safety in numbers, a clear set of instructions, the comfort of a recognizable path. But a factory has one purpose: to produce more of the same.
If you are making Fords, more of the same is good. But
RDHGQE3C
journalArticle
Neel Nanda
Last updated Sept 2 2025 • Note - if you want to pursue a career in this kind of research, apply to my MATS stream! Due Dec 23 …
MD92Z73A
conferencePaper
Andrew Boutros,
Eriko Nurvitadhi,
Rui Ma,
Sergey Gribok,
Zhipeng Zhao,
James C. Hoe,
Vaughn Betz,
Martin Langhammer
The growing importance and compute demands of artificial intelligence (AI) have led to the emergence of domainoptimized hardware platforms. For example, Nvidia GPUs introduced specialized tensor cores for matrix operations to speed up deep learning (DL) computation, resulting in very high peak throughput up to 130 int8 TOPS in the T4 GPU. Recently, Intel introduced its first AI-optimized 14nm FPGA, the Stratix 10 NX, with in-fabric AI tensor blocks that offer estimated peak performance up to 143 int8 TOPS, comparable to 12nm GPUs. However, what matters in practice is not the peak performance but the actual achievable performance on target workloads. This depends mainly on the utilization of the tensor units, and the system-level overheads to send data to/from the accelerator.
E37N55L2
journalArticle
Pouya Kananian,
Arnesh Sujanani,
Seyed Majid Zahedi
We study the fair and truthful allocation of m divisible public items among n agents, each with distinct preferences for the items. To aggregate agents’ preferences fairly, we focus on finding a core solution. For divisible items, a core solution always exists and can be calculated by maximizing the Nash welfare objective. However, such a solution is easily manipulated; agents might have incentives to misreport their preferences. To mitigate this, the current state-of-the-art finds an approximate core solution with high probability while ensuring approximate truthfulness. However, this approach has two main limitations. First, due to several approximations, the approximation error in the core could grow with n, resulting in a non-asymptotic core solution. This limitation is particularly significant as public-good allocation mechanisms are frequently applied in scenarios involving a large number of agents, such as the allocation of public tax funds for municipal projects. Second, implementing the current approach for practical applications proves to be a highly nontrivial task. To address these limitations, we introduce PPGA, a (differentially) Private Public-Good Allocation algorithm, and show that it attains asymptotic truthfulness and finds an asymptotic core solution with high probability. Additionally, to demonstrate the practical applicability of our algorithm, we implement PPGA and empirically study its properties using municipal participatory budgeting data.
V3H2JTRM
preprint
Jiashu Zhang,
Zihan Pan,
Molly,
Xu,
Khuzaima Daudjee,
Sihang Liu
The occurrence of bubbles in pipeline parallelism is an inherent limitation that can account for more than 40% of the large language model (LLM) training time and is one of the main reasons for the underutilization of GPU resources in LLM training. Harvesting these bubbles for GPU side tasks can increase resource utilization and reduce training costs but comes with challenges. First, because bubbles are discontinuous with various shapes, programming side tasks becomes difficult while requiring excessive engineering effort. Second, a side task can compete with pipeline training for GPU resources and incur significant overhead. To address these challenges, we propose FreeRide, a system designed to harvest bubbles in pipeline parallelism for side tasks. FreeRide provides programmers with interfaces to implement side tasks easily, manages bubbles and side tasks during pipeline training, and controls access to GPU resources by side tasks to reduce overhead. We demonstrate that FreeRide achieves 7.8% average cost savings with a negligible overhead of about 1% in training LLMs while serving model training, graph analytics, and image processing side tasks.
9XLB5W5V
preprint
Desen Sun,
Shuncheng Jie,
Sihang Liu
Diffusion models are a powerful class of generative models that produce images and other content from user prompts, but they are computationally intensive. To mitigate this cost, recent academic and industry work has adopted approximate caching, which reuses intermediate states from similar prompts in a cache. While efficient, this optimization introduces new security risks by breaking isolation among users. This paper provides a comprehensive assessment of the security vulnerabilities introduced by approximate caching. First, we demonstrate a remote covert channel established with the approximate cache, where a sender injects prompts with special keywords into the cache system and a receiver can recover that even after days, to exchange information. Second, we introduce a prompt stealing attack using the approximate cache, where an attacker can recover existing cached prompts from hits. Finally, we introduce a poisoning attack that embeds the attacker's logos into the previously stolen prompt, leading to unexpected logo rendering for the requests that hit the poisoned cache prompts. These attacks are all performed remotely through the serving system, demonstrating severe security vulnerabilities in approximate caching. The code for this work is available.
7FSRP4ID
conferencePaper
Yaoyao Ding,
Bohan Hou,
Xiao Zhang,
Allan Lin,
Tianqi Chen,
Cody Hao Yu,
Yida Wang,
Gennady Pekhimenko
YV7VS5YG
webpage
Julia Evans
Get your work recognized: write a brag document
GYJ86NLY
preprint
Mohammad Asadi,
Jack W. O'Sullivan,
Fang Cao,
Tahoura Nedaee,
Kamyar Rajabalifardi,
Fei-Fei Li,
Ehsan Adeli,
Euan Ashley
Multimodal AI systems have achieved remarkable performance across a broad range of real-world tasks, yet the mechanisms underlying visual–language reasoning remain surprisingly poorly understood. We report three findings that challenge prevailing assumptions about how these systems process and integrate visual information. First, Frontier models readily generate detailed image descriptions and elaborate reasoning traces, including pathology-biased clinical findings, for images never provided; we term this phenomenon mirage reasoning. Second, without any image input, models also attain strikingly high scores across general and medical multimodal benchmarks, bringing into question their utility and design. In the most extreme case, our model achieved the top rank on a standard chest Xray question-answering benchmark without access to any images. Third, when models were explicitly instructed to guess answers without image access, rather than being implicitly prompted to assume images were present, performance declined markedly. Explicit guessing appears to engage a more conservative response regime, in contrast to the mirage regime in which models behave as though images have been provided. These findings expose fundamental vulnerabilities in how visual–language models reason and are evaluated, pointing to an urgent need for private benchmarks that eliminate textual cues enabling non-visual inference, particularly in medical contexts where miscalibrated AI carries the greatest consequence. We introduce B-Clean as a principled solution for fair, vision-grounded evaluation of multimodal AI systems.
4IRKFE62
webpage
9S7RH8QG
journalArticle
Curry-Howard Correspondence
Giselle Reis