Field of attention
Paste a link, paper, PDF, or a blob of text. Read it one band at a time — everything outside the band softly blurs, and anything you underline stays sharp and is remembered, so the page distills down to what you kept.
Nothing opened or saved yet.
METR Time Horizon
https://metr.org/time-horizons/The Power of Intelligence
https://www.youtube.com/watch?v=fkWl1Xu-PYsEpoch AI Trends Page
https://epochai.org/trendsCan AI scaling continue through 2030?
https://epoch.ai/blog/can-ai-scaling-continue-through-2030What will AI look like in 2030?
https://epoch.ai/blog/what-will-ai-look-like-in-2030From AGI to Superintelligence
https://situational-awareness.ai/from-agi-to-superintelligence/Inference Scaling and the Log-x Chart
https://forum.effectivealtruism.org/posts/zNymXezwySidkeRun/inference-scaling-and-the-log-x-chartSections about each specific constraint in Can AI scaling continue through 2030?
https://epoch.ai/blog/can-ai-scaling-continue-through-2030Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/Future ML Systems will be Qualitatively Different
https://bounded-regret.ghost.io/future-ml-systems-will-be-qualitatively-different/Training Compute-Optimal Large Language Models
https://arxiv.org/pdf/2203.15556.pdfBiological Anchors: A Trick That Might Or Might Not Work
https://astralcodexten.substack.com/p/biological-anchors-a-trick-that-mightNeural scaling laws and GPT-3
https://www.youtube.com/watch?v=4o4d8F8kJLMExplaining Neural Scaling Laws
https://arxiv.org/abs/2102.06701The Hanson-Yudkowsky AI-Foom Debate
https://intelligence.org/files/AIFoomDebate.pdfSpecification gaming: the flip side of AI ingenuity
https://www.deepmind.com/blog/specification-gaming-the-flip-side-of-ai-ingenuityAligning language models to follow instructions
https://openai.com/blog/instruction-following/Language Models Learn to Mislead Humans via RLHF
https://arxiv.org/pdf/2409.12822Large Language Models can Strategically Deceive their Users when Put Under Pressure
https://arxiv.org/pdf/2311.07590Deep RL from human preferences
https://openai.com/blog/deep-reinforcement-learning-from-human-preferences/Playing the Training Game
https://www.planned-obsolescence.org/the-training-game/Natural Emergent Misalignment from Reward Hacking in Production RL
https://arxiv.org/abs/2511.18397Constitutional AI: Harmlessness from AI Feedback
https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedbackOpen Problems and Fundamental Limitations of RLHF
https://arxiv.org/abs/2307.15217Sycophancy to subterfuge: Investigating reward tampering in language models
https://www.anthropic.com/research/reward-tamperingScaling Laws for Reward Model Overoptimization
https://arxiv.org/abs/2210.10760Concrete Problems in AI Safety
https://arxiv.org/abs/1606.06565Scalable agent alignment via reward modeling
https://arxiv.org/abs/1811.07871The OTHER AI Alignment Problem: Mesa-Optimizers and Inner Alignment
https://www.youtube.com/watch?v=bJLcIBixGj8Alignment faking in large language models
https://arxiv.org/abs/2412.14093Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
https://arxiv.org/abs/2401.05566From shortcuts to sabotage: natural emergent misalignment from reward hacking
https://arxiv.org/abs/2511.18397ML systems will have weird failure modes
https://bounded-regret.ghost.io/ml-systems-will-have-weird-failure-modes-2/Sycophancy to subterfuge: Investigating reward tampering in language models
https://www.anthropic.com/research/reward-tamperingEmergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
https://arxiv.org/abs/2502.07209Frontier Models are Capable of In-Context Scheming
https://www.apolloresearch.ai/research/frontier-models-are-capable-of-in-context-scheming/Distillation of "How likely is deceptive alignment?"
https://www.lesswrong.com/posts/XKraEJrQRfzbCtzKN/distillation-of-how-likely-is-deceptive-alignmentWhy alignment could be hard with modern deep learning
https://www.alignmentforum.org/posts/CoZhXrhpQxpy88Sq8/why-alignment-could-be-hard-with-modern-deep-learningGoal misgeneralization: why correct specifications aren't enough for correct goals
https://arxiv.org/abs/2210.01790The alignment problem from a deep learning perspective
https://arxiv.org/abs/2209.00626Goal Misgeneralization in Deep Reinforcement Learning
https://arxiv.org/abs/2105.14111Optimal policies tend to seek power
https://arxiv.org/abs/1906.01820Language Models as Agent Models
https://arxiv.org/abs/2212.01681What is the orthogonality thesis? (AISafety.info)
https://aisafety.info/questions/6568/What-is-the-orthogonality-thesisWhy Would AI Want to do Bad Things? Instrumental Convergence (Robert Miles, 2018)
https://www.youtube.com/watch?v=ZeecOKBus3QStatement on AI Risk (Center for AI Safety, 2023)
https://safe.ai/work/statement-on-ai-riskAn Overview of Catastrophic AI Risks (Center for AI Safety, 2023)
https://arxiv.org/abs/2306.12001We're Not Ready for Superintelligence (AI in Context, 2025)
https://www.youtube.com/watch?v=4kDPxbS6ofwIntelligence and Stupidity: The Orthogonality Thesis (Robert Miles, 2018)
https://youtu.be/hEUO6pjwFOoAI 2027 (AI Futures Project, 2025)
https://ai-2027.com/The Problem (MIRI, 2025)
https://intelligence.org/the-problem/The Basic AI Drives (Omohundro, 2008)
https://selfawaresystems.com/wp-content/uploads/2008/01/ai_drives_final.pdfGradual Disempowerment (Kulveit et al., 2025)
https://arxiv.org/abs/2501.16946The Superintelligent Will: Motivation And Instrumental Rationality in Advanced Artificial Agents (Nick Bostrom, 2012)
https://nickbostrom.com/superintelligentwill.pdfExistential Risk from Power-Seeking AI (Joe Carlsmith, 2022)
https://arxiv.org/abs/2206.13353Instrumental convergence (Eliezer Yudkowsky, 2025)
https://arbital.com/p/instrumental_convergence/Why Would AI "Aim" To Defeat Humanity? (Cold Takes, 2022)
https://www.cold-takes.com/why-would-ai-aim-to-defeat-humanity/AI Could Defeat All Of Us Combined (Cold Takes, 2022)
https://www.cold-takes.com/ai-could-defeat-all-of-us-combined/Two types of AI existential risk: decisive and accumulative (Atoosa Kasirzadeh, 2025)
https://arxiv.org/abs/2401.07836International AI Safety Report 2026
https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026The Vulnerable World Hypothesis (Nick Bostrom, 2019)
https://nickbostrom.com/papers/vulnerable.pdfAI-enabled coups: a small group could use AI to seize power (Forethought, 2025)
https://www.forethought.org/research/ai-enabled-coups-how-a-small-group-could-use-ai-to-seize-powerImpact of AI on cyber threat from now to 2027 (NCSC, 2025)
https://www.ncsc.gov.uk/report/impact-ai-cyber-threat-now-2027How AI Threatens Democracy (Kreps & Kriner, 2023)
https://www.journalofdemocracy.org/articles/how-ai-threatens-democracy/Can Democracy Survive the Disruptive Power of AI? (Carnegie Endowment for International Peace, 2024)
https://carnegieendowment.org/research/2024/12/can-democracy-survive-the-disruptive-power-of-aiThe Authoritarian Risks of AI Surveillance (Lawfare, 2025)
https://www.lawfaremedia.org/article/the-authoritarian-risks-of-ai-surveillanceThe Operational Risks of AI in Large-Scale Biological Attacks (RAND, 2024)
https://www.rand.org/pubs/research_reports/RRA2977-2.htmlCan we scale human feedback for complex AI tasks? An intro to scalable oversight. (BlueDot Impact, 2025)
https://bluedot.org/blog/scalable-oversight-introChain of Thought Monitorability: A New and Fragile Opportunity for AI Safety (2025)
https://arxiv.org/abs/2507.11473The case for ensuring that powerful AIs are controlled (Redwood Research, 2024)
https://www.lesswrong.com/posts/kcKrE9mzEHrdqtDpE/the-case-for-ensuring-that-powerful-ais-are-controlledAI Control: Improving Safety Despite Intentional Subversion (Redwood Research, 2024)
https://arxiv.org/abs/2312.06942Ctrl-Z: Controlling AI Agents via Resampling (Redwood Research, 2025)
https://arxiv.org/abs/2504.10374An Overview of Control Measures (Redwood, 2025)
https://www.lesswrong.com/posts/G8WwLmcGFa4H6Ld9d/an-overview-of-control-measuresSimple probes can catch Sleeper Agents (Anthropic, 2024)
https://www.anthropic.com/research/probes-catch-sleeper-agentsDebating with More Persuasive LLMs Leads to More Truthful Answers (Khan et al., 2024)
https://arxiv.org/abs/2402.06782Combining W2SG with Scalable Oversight (Leike, 2023)
https://aligned.substack.com/p/combining-w2sg-with-scalable-oversightHumans consulting HCH (Christiano, 2016)
https://ai-alignment.com/humans-consulting-hch-f893f6051455Measuring Progress on Scalable Oversight for Large Language Models (Anthropic, 2022)
https://arxiv.org/abs/2211.03540Debate update: obfuscated arguments problem (OpenAI, 2020)
https://www.alignmentforum.org/posts/PJLABqQ962hZEqhdB/debate-update-obfuscated-arguments-problemAI Safety via Debate (OpenAI, 2018)
https://arxiv.org/abs/1805.00899Supervising Strong Learners by Amplifying Weak Experts (OpenAI, 2018)
https://arxiv.org/abs/1810.08575How to prevent collusion when using untrusted models to monitor each other (Redwood, 2024)
https://www.lesswrong.com/posts/GCqoks9eZDfpL8L3Q/how-to-prevent-collusion-when-using-untrusted-models-toThoughts on the conservative assumptions in AI control (Redwood, 2025)
https://www.lesswrong.com/posts/rHyPtvfnvWeMv7Lkb/thoughts-on-the-conservative-assumptions-in-ai-controlA Sketch of an AI Control Safety Case (UK AISI / Redwood, 2025)
https://arxiv.org/abs/2501.17315Subversion Strategy Eval (Mallen, Griffin, Wagner, Abate, Shlegeris, 2024)
https://arxiv.org/abs/2412.12480An Overview of Areas of Control Work (Redwood, Apr 2025)
https://www.lesswrong.com/posts/Eeo9NrXeotWuHCgQW/an-overview-of-areas-of-control-workFour Places Where You Can Put LLM Monitoring (Redwood, 2025)
https://www.alignmentforum.org/posts/AmcEyFErJc9TQ5ySF/four-places-where-you-can-put-llm-monitoringMisalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hacking (Redwood, May 2025)
https://www.lesswrong.com/posts/TeTegzR8X5CuKgMc3/misalignment-and-strategic-underperformance-an-analysis-ofHidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs (LASR Labs, 2024)
https://arxiv.org/abs/2410.03768Detecting Strategic Deception Using Linear Probes (Apollo Research, 2025)
https://arxiv.org/abs/2502.03407Debate Helps Supervise Unreliable Experts (Michael et al., 2023)
https://arxiv.org/abs/2311.08702How to keep improving when you're better than any teacher — IDA (Robert Miles, 2019)
https://www.youtube.com/watch?v=v9M2Ho9I9QoThe Stop Button Problem and corrigibility (Robert Miles / Computerphile, 2017)
https://www.youtube.com/watch?v=3TYT1QfdfsMRob Miles YouTube explainer of the AI Control paper (Robert Miles, 2024)
https://www.youtube.com/watch?v=0pgEMWy70QkAXRP Episode 27: AI Control with Buck Shlegeris and Ryan Greenblatt (2024)
https://axrp.net/episode/2024/04/11/episode-27-ai-control-buck-shlegeris-ryan-greenblatt.html80,000 Hours: Buck Shlegeris on controlling AI that wants to take over (2024)
https://80000hours.org/podcast/episodes/buck-shlegeris-ai-control-scheming/80,000 Hours: Ryan Greenblatt on AI R&D automation and misaligned takeover (2025)
https://80000hours.org/podcast/episodes/ryan-greenblatt-ai-automation-sabotage-takeover/Reasoning Models Don't Always Say What They Think (Chen et al., Anthropic/OpenAI, 2025)
https://www.anthropic.com/research/reasoning-models-dont-say-thinkMonitoring Monitorability (Guan, Wang, Carroll et al., OpenAI, 2025)
https://arxiv.org/abs/2512.18311Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity (Meek et al., 2025)
https://arxiv.org/abs/2510.27378Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps (EMNLP 2025)
https://arxiv.org/abs/2502.14829Language Models Don't Always Say What They Think (Turpin et al., 2023)
https://arxiv.org/abs/2305.04388Measuring Faithfulness in Chain-of-Thought Reasoning (Lanham et al., 2023)
https://arxiv.org/abs/2307.13702Chain-of-Thought Reasoning in the Wild: Faithfulness (Arcuschin et al., 2025)
https://arxiv.org/abs/2503.08679When Chain of Thought Is Necessary, Language Models Struggle to Evade Monitors (Emmons et al., 2025)
https://arxiv.org/abs/2507.05246CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring (Arnav et al., 2025)
https://arxiv.org/abs/2505.23575LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring (Li, Phuong, Siegel, 2025)
https://arxiv.org/abs/2508.00943Stress Testing Deliberative Alignment for Anti-Scheming Training (Schoen et al., Apollo Research, 2025)
https://arxiv.org/abs/2509.15541Auditing Language Models for Hidden Objectives (Marks et al., Anthropic, 2025)
https://www.anthropic.com/research/auditing-hidden-objectivesPreventing Language Models From Hiding Their Reasoning (Roger & Greenblatt, Redwood Research, 2023)
https://arxiv.org/abs/2310.18512Situational Awareness: The Decade Ahead
https://situational-awareness.ai/wp-content/uploads/2024/06/situationalawareness.pdfLearning to reason with LLMs
https://openai.com/index/learning-to-reason-with-llms/Combining Deep Reinforcement Learning and Search for Imperfect-Information Games
https://arxiv.org/abs/2007.13544Human-Level Performance in No-Press Diplomacy via Equilibrium Search
https://arxiv.org/abs/2010.02923Attention Is All You Need
https://arxiv.org/abs/1706.03762Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
https://arxiv.org/abs/1701.06538Fast Transformer Decoding: One Write-Head is All You Need
https://arxiv.org/abs/1911.02150Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
https://arxiv.org/abs/2101.03961Machine Psychology
https://arxiv.org/abs/2303.13988Evaluating Large Language Models with Psychometrics
https://arxiv.org/abs/2406.17675Towards Understanding Sycophancy in Language Models
https://arxiv.org/abs/2310.13548"Dark Triad" Model Organisms of Misalignment
https://arxiv.org/abs/2603.06816Persona vectors: Monitoring and controlling character traits in language models
https://www.anthropic.com/research/persona-vectorsMapping the Mind of a Large Language Model
https://www.anthropic.com/research/mapping-mind-language-modelTracing the thoughts of a large language model
https://www.anthropic.com/research/tracing-thoughts-language-modelTowards Monosemanticity: Decomposing Language Models With Dictionary Learning
https://www.anthropic.com/research/towards-monosemanticity-decomposing-language-models-with-dictionary-learningDiscovering Language Model Behaviors with Model-Written Evaluations
https://www.anthropic.com/research/discovering-language-model-behaviors-with-model-written-evaluationsAlignment faking in large language models
https://www.anthropic.com/research/alignment-fakingSleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
https://www.anthropic.com/research/sleeper-agents-training-deceptive-llms-that-persist-through-safety-trainingConstitutional AI: Harmlessness from AI Feedback
https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedbackSycophancy in GPT-4o: what happened and what we're doing about it
https://openai.com/index/sycophancy-in-gpt-4o/Sharing the latest Model Spec
https://openai.com/index/sharing-the-latest-model-spec/Detecting misbehavior in frontier reasoning models
https://openai.com/index/chain-of-thought-monitoring/Deliberative alignment: reasoning enables safer language models
https://openai.com/index/deliberative-alignment/Weak-to-strong generalization
https://openai.com/index/weak-to-strong-generalization/Language models can explain neurons in language models
https://openai.com/index/language-models-can-explain-neurons-in-language-models/Extracting concepts from GPT-4
https://openai.com/index/extracting-concepts-from-gpt-4/Multimodal neurons in artificial neural networks
https://openai.com/index/multimodal-neurons/LIMA: Less Is More for Alignment
https://arxiv.org/abs/2305.11206Shepherd: A Critic for Language Model Generation
https://arxiv.org/abs/2308.04592Meta-Rewarding Language Models
https://arxiv.org/abs/2407.19594Toolformer: Language Models Can Teach Themselves to Use Tools
https://ai.meta.com/research/publications/toolformer-language-models-can-teach-themselves-to-use-tools/Specification gaming: the flip side of AI ingenuity
https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/Goal Misgeneralization: Why Correct Specifications Aren't Enough For Correct Goals
https://arxiv.org/abs/2210.01790Evaluating Frontier Models for Dangerous Capabilities
https://arxiv.org/abs/2403.13793Evaluating Frontier Models for Stealth and Situational Awareness
https://arxiv.org/abs/2505.01420