Zotero Catalogue
A public reading list of papers, books, videos, and other resources. The inclusion of a resource on this catalogue is NOT an endorsement of anything contained within, and in most cases the resources has not been read by me at the time of saving.
1157 items · showing 701–750 · page 15 of 24 Sort: Newest Oldest Title A–Z Title Z–A
X42MRP5L
preprint
Luke Darlow,
Ciaran Regan,
Sebastian Risi,
Jeffrey Seely,
Llion Jones
Biological brains demonstrate complex neural activity, where neural dynamics are critical to how brains process information. Most artificial neural networks ignore the complexity of individual neurons . We challenge that paradigm. By incorporating neuron-level processing and synchronization, we reintroduce neural timing as a foundational element. We present the Continuous Thought Machine (CTM), a model designed to leverage neural dynamics as its core representation. The CTM has two innovations: (1) neuron-level temporal processing, where each neuron uses unique weight parameters to process incoming histories; and (2) neural synchronization as a latent representation. The CTM aims to strike a balance between neuron abstractions and biological realism. It operates at a level of abstraction that effectively captures essential temporal dynamics while remaining computationally tractable. We demonstrate the CTM’s performance and versatility across a range of tasks, including solving 2D mazes, ImageNet1K classification, parity computation, and more. Beyond displaying rich internal representations and offering a natural avenue for interpretation owing to its internal process, the CTM is able to perform tasks that require complex sequential reasoning. The CTM can also leverage adaptive compute, where it can stop earlier for simpler tasks, or keep computing when faced with more challenging instances. The goal of this work is to share the CTM and its associated innovations, rather than pushing for new state-of-the-art results. To that end, we believe the CTM represents a significant step toward developing more biologically plausible and powerful artificial intelligence systems. We provide an accompanying interactive online demonstration and an extended technical report.
HBX85V6G
webpage
Sakana AI
新しいBusiness Intelligenceへ:Ultra Deep Researchアシスタント「Sakana Marlin」βテスト開始
RAZRJ4GH
blogPost
Mattan Griffel
5 rules for good email etiquette
LJF8XL3T
webpage
RGDWELX2
webpage
Parenthetically Speaking: Articles by Shriram Krishnamurthi
IAL827EV
webpage
Cohere Labs Community
A community lead reflects on three years of learning, research, and building programs inside the Cohere Labs Open Science Community.
7UVC593T
webpage
Patrick Lewis,
Ethan Perez,
Aleksandra Piktus,
Fabio Petroni,
Vladimir Karpukhin,
Naman Goyal,
Heinrich Küttler,
Mike Lewis
et al.
Large pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks. However, their ability to access and precisely manipulate knowledge is still limited, and hence on knowledge-intensive tasks, their performance lags behind task-specific architectures. Additionally, providing provenance for their decisions and updating their world knowledge remain open research problems. Pre-trained models with a differentiable access mechanism to explicit non-parametric memory can overcome this issue, but have so far been only investigated for extractive downstream tasks. We explore a general-purpose fine-tuning recipe for retrieval-augmented generation (RAG) -- models which combine pre-trained parametric and non-parametric memory for language generation. We introduce RAG models where the parametric memory is a pre-trained seq2seq model and the non-parametric memory is a dense vector index of Wikipedia, accessed with a pre-trained neural retriever. We compare two RAG formulations, one which conditions on the same retrieved passages across the whole generated sequence, the other can use different passages per token. We fine-tune and evaluate our models on a wide range of knowledge-intensive NLP tasks and set the state-of-the-art on three open domain QA tasks, outperforming parametric seq2seq models and task-specific retrieve-and-extract architectures. For language generation tasks, we find that RAG models generate more specific, diverse and factual language than a state-of-the-art parametric-only seq2seq baseline.
7W53YSF6
webpage
What is the reason for proliferation of DSLs in the last year?
86K27KWP
webpage
Form and track positive lasting habits built with 💙 by me - powered by org 🦄 Why Keeping habits accessible and trackable has helped me form good hab...
5C3D28RH
journalArticle
Journal of Natural Science and Exploration
Ekagrata Bahadur,
Amrit Nath Thulal
A Natural language processing (NLP) has increased the interest in genetic algorithm (GA) due to their skills in solving complex optimization problems with extensive research on the use of genetic algorithms in NLP projects has been presented in this paper. First, we present the basic concepts behind genetic algorithms and their relevance to natural language processing. Then, we explore various applications of natural language processing (NLP) that use genetic algorithms, including text classification, sentiment analysis, machine translation, summarization, and question-answering systems. We examine the advantages and disadvantages of genetic algorithm applications in natural language processing by comparing their performance with traditional and modern approaches and discuss the factors influencing their effectiveness. Furthermore, we explore recent advancements, modifications, and hybridizations of Genetic Algorithms tailored to NLP tasks. Finally, we discuss the challenges and future directions in leveraging Genetic Algorithms for enhancing NLP technologies.
B7XFUCZM
webpage
Confluent is building the foundational platform for data in motion so any organization can innovate and win in a digital-first world.
CT4L46N4
webpage
A few days ago, I posted about a personal project that I've been working on for the last few weeks: Mr. Chatterbox, a chatbot trained from scratch on Victorian-era literature. I have to admit, I was totally blown away (and a little frightened!) by the reception. Whenever I post about
G3ZFHKRY
webpage
Media over QUIC: There are ways to do voice AI without being traumatized by WebRTC.
UWSSKBJ6
webpage
COS568 Systems and Machine Learning (Spring 2025) Programming Assignments Network pruning Distributed training of a language model Project Learned index Tentative Syllabus Dates Presenters Topics & Main Papers Related Papers Events 1/31 Kai Li Dr. Jeff Dean & Dr. Amin Vahdat (Goo...
3HT5CZX7
webpage
How the rising star of podcasting eschews the breadth v. depth dilemma and is quickly becoming known as ‘the new Lex Fridman.’
ADWIMTN9
newspaperArticle
Katya Ungerman
One of the oldest and most durable features of human experience is re-emerging.
DW9GGNSS
webpage
We launched the Datasette Cloud blog today. The Datasette Cloud site itself is a Django app - it uses Django and PostgreSQL to manage accounts, teams and soon billing and payments, then launches dedicated containers running Datasette for each customer.
Y5IKS6HK
webpage
Nextpad++ feels like a fever dream. Like what Mac apps would be if the Nazis had won WWII.
5DRBECLG
webpage
It’s not even a feature. It’s just technology.
QUWPVPZS
webpage
Simon Willison
François Chollet is the co-founder of the ARC Prize and had advanced access to today's o3 results. His article here is the most insightful coverage I've seen of o3, going …
W3KQRHHN
webpage
Simon Willison
You should start a blog. Having your own little corner of the internet is good for the soul! But what should you write about? It’s easy to get hung up …
NBI5SJ4C
webpage
POSSE is an abbreviation for Publish (on your) Own Site, Syndicate Elsewhere, the practice of posting content on your own site first, then publishing copies or sharing links to third parties (like social media silos) with original post links to provide viewers a path to directly interacting with your content.
RIKWAVLC
webpage
Simon Willison
I started running a basic link blog on this domain back in November 2003—publishing links (which I called “blogmarks”) with a title, URL, short snippet of commentary and a “via” …
469NQS3D
webpage
Jason Koebler ·
AI writing is impossible to avoid, is making everything sound the same, and is driving us crazy.
PEJMAT5S
blogPost
Zohar Atkins
Walter Benjamin and Rabbi Levi Yitzchak of Berditchev on Translation
ZKJLDT7I
webpage
Notes on applied AI engineering, machine learning, and data science.
3H73EWEN
blogPost
Malmesbury
Recommended soundtrack for this post
KWTDAYJB
preprint
Haoyi Zhu,
Haozhe Liu,
Yuyang Zhao,
Tian Ye,
Junsong Chen,
Jincheng Yu,
Tong He,
Song Han
et al.
We introduce SANA-WM, an efficient 2.6B-parameter open-source world model natively trained for one-minute generation, synthesizing high-fidelity, 720p, minute-scale videos with precise camera control. SANA-WM achieves visual quality comparable to large-scale industrial baselines such as LingBot-World and HY-WorldPlay, while significantly improving efficiency. Four core designs drive our architecture: (1) Hybrid Linear Attention combines frame-wise Gated DeltaNet (GDN) with softmax attention for memory-efficient long-context modeling. (2) Dual-Branch Camera Control ensures precise 6-DoF trajectory adherence. (3) Two-Stage Generation Pipeline applies a long-video refiner to stage-1 outputs, improving quality and consistency across sequences. (4) Robust Annotation Pipeline extracts accurate metric-scale 6-DoF camera poses from public videos to yield high-quality, spatiotemporally consistent action labels. Driven by these designs, SANA-WMdemonstrates remarkable efficiency across data, training compute, and inference hardware: it uses only $\sim$213K public video clips with metric-scale pose supervision, completes training in 15 days on 64 H100s, and generates each 60s clip on a single GPU; its distilled variant can be deployed on a single RTX 5090 with NVFP4 quantization to denoise a 60s 720p clip in 34s. On our one-minute world-model benchmark, SANA-WM demonstrates stronger action-following accuracy than prior open-source baselines and achieves comparable visual quality at $36\times$ higher throughput for scalable world modeling.
PYV9I8V9
preprint
Hao Tang,
Darren Key,
Kevin Ellis
We give a model-based agent that builds a Python program representing its knowledge of the world based on its interactions with the environment. The world model tries to explain its interactions, while also being optimistic about what reward it can achieve. We define this optimism as a logical constraint between a program and a planner. We study our agent on gridworlds, and on task planning, finding our approach is more sample-efficient compared to deep RL, more compute-efficient compared to ReAct-style agents, and that it can transfer its knowledge across environments by editing its code.
HK2DKCS7
forumPost
ConradIsMyDaddy
J5KZDL4H
webpage
3 track album
6WQ4QICZ
journalArticle
Junhyuk Oh,
Gregory Farquhar,
Iurii Kemaev,
Dan A. Calian,
Matteo Hessel,
Luisa Zintgraf,
Satinder Singh,
Hado van Hasselt
et al.
Humans and other animals use powerful reinforcement learning (RL) mechanisms that have been discovered by evolution over many generations of trial and error. By contrast, artificial agents typically learn using handcrafted learning rules. Despite decades of interest, the goal of autonomously discovering powerful RL algorithms has proven to be elusive1–6. Here we show that it is possible for machines to discover a state-of-the-art RL rule that outperforms manually designed rules. This was achieved by meta-learning from the cumulative experiences of a population of agents across a large number of complex environments. Specifically, our method discovers the RL rule by which the agent’s policy and predictions are updated. In our large-scale experiments, the discovered rule surpassed all existing rules on the well-established Atari benchmark and outperformed a number of state-of-the-art RL algorithms on challenging benchmarks that it had not seen during discovery. Our findings suggest that the RL algorithms required for advanced artificial intelligence may soon be automatically discovered from the experiences of agents, rather than manually designed.
JESHI4F9
webpage
Thinking Machines Lab
Interaction models move beyond turn-based AI interfaces by handling multimodal, real-time collaboration natively across audio, video, and text.
4PLTSZBD
book
Stoner
John Williams
UVLHGE6F
forumPost
Adventurous-Math-322