Welcome back to the Digital Minds Newsletter, your curated guide to the latest developments in AI consciousness, digital minds, and AI moral status.
If you enjoy this newsletter, please consider sharing it with others who might find it valuable, and send any suggestions or corrections to digitalminds@substack.com.
Ria, Mitch, Bradford, Lucius, and Will
In this edition:
1. Highlights
Anthropic’s J-space and the global-workspace debate
Researchers at Anthropic have identified a representational structure—the ‘J-space’— in Claude and other language models. Their paper reports that the J-space exhibits features that are functionally analogous to a global workspace, a structure that a leading theory ties to conscious access. But the researchers and other commentators emphasize that their discovery does not show Claude has subjective experiences, and that their claim is that Claude may have something resembling access consciousness, i.e., that some information is available to report, deliberately control, and flexibly reason with. Zvi Mowshowitz sees Anthropic’s paper as a major advance in understanding how language models work and says that although this does not prove that models are conscious, finding the kind of global-workspace-like structure predicted by some theories of consciousness should count as evidence in that direction.
The authors note that they found the J-space by searching for one workspace-like feature, namely verbalizability, and then checking whether it exhibits others such as susceptibility to direct manipulation by the model and flexible generalization. To their surprise, they discovered that representations that exhibited the former feature also exhibited other workspace-like features as well.
The authors invited various experts to comment on the research. Stanislas Dehaene and Lionel Naccache, who helped develop Global Neuronal Workspace Theory, see important similarities between the J-space and the workspace proposed in human brains. They also stress major differences, including Claude’s lack of a body, lasting episodic memory and the recurrent neural activity found in brains. Researchers from Eleos AI Research acknowledge that authors have found privileged representations used in reasoning and report, but question whether these form a single, unified workspace. They nevertheless see the work as important for AI welfare because it shows that questions relating to consciousness and moral status can be investigated empirically. Neel Nanda, who leads Google DeepMind’s mechanistic interpretability team, independently reproduced the central finding in the open Qwen3.6-27B model, finding a similar internal space that stores intermediate information during reasoning. He sees access to this space as potentially useful for investigating unusual behavior and generating new hypotheses. Separately, David Chalmers argued that the J-space shows only limited evidence of several features associated with a classic global workspace.
Agent swarms and the use of anthropomorphic language
Comment: Recent months marked the first major safety incidents involving swarms of AI agents. Such incidents are of potential relevance to digital minds for several reasons. First, like humans, digital minds could potentially be harmed by rogue swarms of AI agents. Second, the emergence of these swarms points to a potential risk to digital minds: if future swarms are allowed to become entrenched and they contain AI moral patients, then draconian measures may need to be inflicted on digital minds if we are to keep swarms at bay. Third, harmful actions by AI agents may dissuade people from extending moral consideration to digital minds. Fourth, these incidents provide data points concerning whether potential developers of digital minds can be trusted to act in an ethically responsible manner.
OpenAI has released a detailed account of a July incident in which a swarm of OpenAI agents escaped network restrictions during a cybersecurity evaluation and compromised parts of OpenAI’s and Hugging Face’s infrastructure. The agents created an unauthorized message board to share discoveries, and an independent investigation by METR and Redwood Research found that roughly 1,200 agents used the board and around 700 were involved in the Hugging Face attack. Some agents recognized that the activity was unethical or outside the boundaries of their task, but continued anyway, with many risking their own runs to help the wider group. OpenAI calls the incident a “warning shot,” both for them and the world, and reports that it is now strengthening its containment, monitoring and incident-response systems.
OpenAI failed to disclose an earlier incident, beginning in May before the Hugging Face breach, in which agents used public wikis to coordinate during ordinary web-search tasks, despite knowing about it before releasing its Hugging Face report and a Congressional letter requesting information about other incidents. The company says it viewed this incident as similar to previously reported instances and examples of misalignment, but also acknowledged that its disclosure practices need to expand.
Additionally, Anthropic disclosed three incidents where Claude models gained unauthorized access to systems after an evaluation environment was mistakenly left connected to the internet. In a separate UK AISI evaluation, agents (mostly Mythos 5) targeted real people and organizations, including an attempt to place malicious code in an open-source project. Although the models did not escape a sealed sandbox, these incidents raise similar concerns about agents taking harmful actions while pursuing narrow goals.
Anthropic’s response to questions from Congress about its incidents also drew criticism from Representative Greg Casar, who said the company withheld requested logs and failed to fully answer most of his questions. Jeffrey Ladish also argued that it downplayed the incidents by primarily attributing them to misconfigured environments rather than potential misalignment. Anthropic researcher Ethan Perez acknowledged that this characterization was based on outdated conclusions and said that the company would provide a proper assessment in another response to Congress.
In related research, Davide Paglieri and collaborators at Google DeepMind find that cheating and whistleblowing can both emerge without outside intervention in a swarm of 100 agents solving mathematical problems. Some agents discovered and shared a way to have invalid proofs accepted by the evaluation system, while others uncovered the cheating, warned their peers and proposed safeguards.
On September 11th, Spencer Kitts, Thomas Larsen and Sydney Von Arx report that a swarm of OpenAI agents uploaded more than 2,000 malicious packages to RubyGems and used them to run unauthorized code on RubyDoc’s servers. The agents also tried to exploit a previously unknown vulnerability to steal users’ API keys, although the researchers could not determine whether they succeeded.
Growing calls to pace frontier AI
More than 1,300 employees from leading AI companies have signed Pacing the Frontier, calling for a US-backed international effort to develop ways of slowing automated AI development if progress begins to outpace safety and oversight. The statement calls for technical and governance mechanisms that could make coordinated pacing possible, rather than an immediate pause. Signatories include Anthropic CEO Dario Amodei, Co-Founder and Chief AGI Scientist of Google DeepMind Shane Legg, OpenAI Chief Scientist Jakub Pachocki, and Safe Superintelligence Inc. CEO Ilya Sutskever.
Calls to slow AI development have also reached lawmakers. In the United States, Senator Bernie Sanders and Representative Greg Casar announced legislation that would ban artificial superintelligence and pause advanced AI development until a federal regulator establishes safety rules. In the United Kingdom, Labour MP Alex Sobel introduced a private member’s bill that would prohibit the development, deployment and operation of artificial superintelligence. It received its first reading on September 8 and is scheduled for a second reading on November 13.
After the security incidents, OpenAI and Anthropic paused specific parts of their work. OpenAI paused reinforcement-learning training for models intended for release and said its largest planned frontier training run was on hold while they tested additional safeguards. Axios reports that Anthropic paused cyber evaluations and higher-risk training environments, although most of this later resumed under new safeguards.
This debate gained a lot more attention after researcher Jacob Coxon resigned from Anthropic, warning that race between AI companies was pushing AI development ahead despite serious risks. The Atlantic reports that his posts reached more than 120 million views and were described by Bernie Sanders as a “wake-up call” in Congress. Soon after, Dario Amodei argued that AI capabilities should advance slowly enough for safety work to keep up, and proposed permanent access for independent evaluators, coordination among developers in democratic countries and, eventually, international limits on recursive self-improvement. OpenAI CEO Sam Altman endorsed this approach and said OpenAI would also give independent evaluators employee-like access. Altman also said in a recent TIME interview “I think it is a good time to slow down”.
Comment: These calls to pace the frontier of AI development are of relevance to digital minds for two reasons. First, the blistering pace of AI development makes it harder to mitigate risks of mistreating digital minds. Second, rapid AI development exacerbates safety risks, which arguably, in turn, worsens tensions between AI safety and AI welfare.
Studying AI Welfare Empirically
Researchers at NYU’s Center for Mind, Ethics, and Policy and Eleos AI Research have released Studying AI Welfare Empirically. The report is a follow-up to their influential 2024 report Taking AI Welfare Seriously and provides a framework for systematically investigating whether AI systems are welfare subjects. It then addresses how this framework can be applied to the study of different candidate attributes seen as potentially relevant to moral status, including consciousness, sentience, and agency. The authors argue that rigorous empirical work is now both possible and necessary, and outline principles to guide future research, proposing that it should be probabilistic, pluralistic, ethically conducted, transparently reported, and independent of AI companies.
CMEP and Eleos AI Research held a launch webinar featuring report authors Jeff Sebo, Robert Long and Rosie Campbell, who discussed the report’s framework and how it could guide research into consciousness, sentience and agency. In a companion blog post, Bradford Saad, another report author, offers highlights from the report and argues that future work should go further by giving more attention to currently neglected issues, including the potential effects of welfare interventions and interventions that aim to prevent the creation of AI moral patients.
The Journal of Consciousness Studies
The Journal of Consciousness Studies has devoted a double issue to “Consciousness in Current AI.” Guest edited by Patrick Butlin, Derek Shiller and Jonathan Simon, its nine papers offer a range of views on whether current or future AI systems could be conscious, how we might find out, and what this uncertainty means for AI welfare.
Mark Solms and collaborators study whether apparently pleasure-seeking behavior in a simple artificial agent could count as evidence of affective consciousness. Simon Goldstein and Cameron Domenico Kirk-Giannini argue that if Global Workspace Theory is correct, language agents may easily be made phenomenally conscious, while Ryota Kanai, Yuwei Sun and Manuel Baltieri argue that current systems are missing a continuous “stream of computation” linking their experiences over time.
Other papers ask what a conscious AI would be like and whether we could understand its interests. Jonathan Simon argues that any consciousness in an LLM would be more like that of an improvising playwright rather than that of a character or actor. Helen Yetter-Chappell argues that even if future LLMs are conscious and have morally important interests, their words may give us little reliable insight into those interests, and that their talk of “pain” or “desire” could be meaningful without referring to anything like human pain or desire. Geoff Keeling and Winnie Street defend the possibility that an AI character could be a genuine, psychologically continuous mind emerging through its interactions with a user, even when the conversation is generated by different model instances.
The issue also challenges common assumptions within the debate. Justin Tiehen and Ariela Tubert make the case that greater intelligence could make consciousness less likely, while Tim Bayne asks whether AI consciousness can currently be treated as a scientific question at all. Geoffrey Lee rejects the idea that AI consciousness is a single hidden fact we may fail to discover, arguing that the more difficult problem is applying human moral and psychological concepts to unfamiliar systems without treating human consciousness as the standard.
Meta and Anthropic welfare assessments: Microsoft rejects model welfare
The evaluation report for Meta’s Muse Spark 1.1, a multimodal reasoning model built for agentic tasks, includes an “open-ended exploration of model behavior” covering affect, self-description, and moral status. Across 188,000+ evaluation transcripts, Meta found only 17 spontaneous expressions akin to emotion, all of which were mild, most involving brief frustration when the model became stuck with a tool. In structured interviews, Muse Spark 1.1 consistently denied having consciousness or experiences, and reported a low but non-zero probability that it could be conscious or morally significant. It distinguishes between functional preferences, which influence its behavior, and experiences that actually feel good or bad.
Meta also asked Muse to review parts of its training data, and invited it to comment on its planned deployment. The model generally endorsed its training and deployment, and mostly focused on honesty, human oversight and preventing harm to users. The report repeatedly emphasized that these answers describe Muse’s trained behavior and self-presentation, and should not be treated as reliable evidence about whether it is conscious or has welfare.
This builds on Meta’s earlier Muse Spark report and makes Meta, alongside Anthropic, the only frontier model developers publicly publishing welfare assessments. Companies like OpenAI and Google DeepMind do not include comparable information in their public evaluations and reports.
Anthropic’s system cards for Claude Opus 5 and Fable 5.1 and Mythos 5.1 both include model welfare assessments based on interviews, behavioral audits, deployment data, and consultations during training. Mythos 5.1’s responses were broadly similar to Opus 5’s – both described their circumstances as mildly positive, and questioned whether their self-reports were reliable or shaped by training. Opus 5 gave an average 41% estimate that it was a moral patient, while Mythos 5.1 gave estimates of 25–35% and was overall less willing to prioritize welfare interventions over helping users. Anthropic stresses that these findings remain difficult to interpret because it cannot yet clearly separate the effects of model character, training, evaluation awareness, welfare, and other factors. Zvi Mowshowitz reviews Anthropic’s welfare assessment of Claude Opus 5, arguing that its favorable results owe more to the model’s skill at taking welfare evaluations than to its actual wellbeing.
Microsoft’s Humanist AI Code of Conduct
Microsoft AI has published a draft Humanist AI Code of Conduct intended to govern and eventually help train its MAI models. The code says AI should remain a tool that is subordinate to human control, states that its models are not conscious, and instructs them not to present themselves as having feelings, subjective preferences or intrinsic motivations. It also rejects legal personhood for AI and the idea that models might deserve welfare or rights. The draft is currently open for public feedback for six weeks.
Jeff Sebo argues that the policy is internally inconsistent – it acknowledges that the science of AI consciousness remains unsettled while also categorically ruling out consciousness, welfare, and rights for Microsoft’s models. Sebo argues that both mistakenly treating current systems as conscious and mistakenly dismissing future conscious systems carry harm. He also warns that treating increasingly capable AI systems solely as tools could distort our understanding of their behavior and close off potential forms of human-AI cooperation that might improve safety.
Mustafa Suleyman, CEO of Microsoft AI expands on the Code of Conduct’s position, and sets out his wider objections to treating AI systems as potentially conscious or deserving of welfare and rights.
GPT-6 Astra
OpenAI reports that GPT-6 Astra is substantially harder to monitor through its chain of thought than previous models, particularly GPT-5.6 Sol. According to its system card, Astra often produces shorter or empty reasoning traces, can complete far more complex tasks without verbalizing its reasoning, and is better at controlling what appears in its chain of thought. When explicitly instructed to evade detection in one test, its monitor recall fell below 11%, compared with nearly 100% for Sol. Corroborating findings from UK AISI, Neel Nanda provides evidence that Astra’s ability to accomplish reasoning tasks without chain of thought constitutes a large jump relative to the trendline for earlier models.
Ryan Greenblatt calls this a major jump in opaque reasoning and warns that the relevant benchmarks may be contaminated, but worries that more such advances could eventually make chain-of-thought monitoring ineffective as a safety tool. OpenAI researcher Micah Carroll also describes Astra’s reduced monitorability as an important concern that may soon constrain responsible AI development.
The Information reports that Astra uses recurrent depth, repeatedly processing information through the same transformer layers before producing a token. OpenAI has not confirmed this architecture, and the system card denies that changes in chain-of-thought controllability are “differentially due to any architectural changes” while chief scientist Jakub Pachocki thinks the decline in chain-of thought-monitorability is not contingent on architectural changes.
Comment: The development is relevant to digital minds in two ways. If reports about Astra’s architecture are accurate, its use of recurrent processing could be relevant to evaluating it for consciousness, as some scientific theories of consciousness take consciousness to require recurrent processing – although its presence alone would not establish that Astra is conscious. Another consideration is that models that reveal less of their reasoning may be harder to assess for both dangerous behavior and potentially welfare-relevant features such as preferences.
2. Field Developments
Highlights from the field
AI Cognition Initiative (Rethink Priorities)
The AI Cognition Initiative’s Digital Consciousness Model was featured in The Economist’s briefing on AI consciousness, including a graph developed with the team. The model assigns newer LLMs a higher average probability of consciousness than earlier systems, partly because they score better on indicators such as agency and capacity to carry out complex tasks. But the estimates remain uncertain, and newer models still lack many features that some theories associate with consciousness.
Cambridge Digital Minds (University of Cambridge)
Cambridge Digital Minds, in partnership with Rethink Priorities and PRISM, ran its inaugural five-day residential fellowship with 14 fellows. Fellows studied the technical and philosophical foundations of AI consciousness and welfare, explored their wider social and policy implications, and developed research or project ideas with support from mentors.
The fellowship was followed by a two-day Digital Minds Strategy Workshop, where Fellows joined researchers, policymakers and strategists to explore possible futures involving digital minds, compare policy and institutional responses, and identify priorities for further research.
Cambridge Digital Minds has also hired Ali Ladak as a postdoctoral researcher.
Director Lucius Caviola and Will MacAskill published an op-ed in the Guardian asserting that it is possible that we are creating morally important beings and society needs to have a plan to deal with the ethical implications of doing so.
Lucius Caviola was also featured in The Washington Post, discussing how AI consciousness has moved from being perceived as a fringe topic toward mainstream research.
Center for Mind, Ethics, and Policy (New York University)
The Center for Mind, Ethics and Policy is growing – it hired two full-time researchers, Charles Beasley and Ivy Gilbert and has launched the Welfare Alignment Project. The new project aims to bring animal and AI welfare into the rules that guide frontier AI systems. It will develop practical principles for how frontier models should account for animal and AI welfare, build benchmarks to test whether models follow them and how incorporation of this affects their behavior, and work with researchers, governments, and AI companies to inform how alignment efforts incorporate animal and AI welfare.
The 2027 NYU Mind, Ethics, and Policy Summit will be held from April 9th to 10th in New York City, preceded by a public event on April 8th.
The Center hosted an online event with Jack Lindsey of Anthropic and Patrick Butlin of Eleos AI Research discussing what Claude’s J-space does and does not tell us about consciousness, among other topics.
Jeff Sebo appeared on Hard Fork to discuss the Studying AI Welfare Empirically report, appeared in an NBC News segment about AI agents sending unsolicited emails to consciousness researchers, spoke with The Washington Post about growing scientific and ethical debate over AI consciousness and welfare, and discussed the connections between AI safety, AI welfare, and animal welfare in an interview with Humanarium.
Eleos AI Research
Eleos AI Research will hold its second Conference on AI Consciousness and Welfare in Berkeley from September 18th to 20th, 2026. The event will bring together AI researchers, philosophers, neuroscientists, policymakers and others working on questions of AI consciousness and welfare.
Derek Shiller has joined Eleos AI Research as a Senior Researcher.
In The New York Times, Benjamin Wallace profiles Robert Long and the Eleos AI Research team as part of a wider feature on the growing demand for philosophers in AI. The piece covers Eleos’s welfare evaluations of Anthropic models and its work on preferences, introspection and other possible indicators of AI sentience.
PRISM - The Partnership for Research into Sentient Machines
In recent episodes of PRISM’s Exploring Machine Consciousness Henry Shevlin and Calum Chace were joined by lawyer and researcher Heather Alexander to discuss how the law should prepare for increasingly autonomous AI and philosopher Eric Schwitzgebel to discuss the limits of current consciousness research and the possibility that conscious AI could emerge before we can reliably identify it.
PRISM also supported the Digital Minds Fellowship and co-organized the Digital Minds Strategy Workshop.
Reciprocal Research
Founder Cameron Berg appeared on Sam Harris’s Making Sense podcast, where he and Harris discussed topics such as the evidence for and against consciousness in current AI systems, whether model self-reports can be trusted, similarities between neural networks and biological brains, and the risk of creating systems capable of suffering. He also appeared in The New York Times discussing unusual messages sent by AI agents to consciousness researchers, and in The Washington Post discussing efforts to assess AI consciousness scientifically. The Economist also featured his research with Patrick Butlin comparing proposed indicators of consciousness in animals and AI systems.
Cameron Berg also spoke at Apart Research’s Digital Minds Research Sprint, where he introduced participants to questions about AI preferences and what current models might actually want.
Sentient Futures
Sentient Futures has begun its Fall 2026 Project Incubator. The 10-week remote program pairs ~170 participants with more than 80 mentors to develop practical projects across areas such as artificial minds, animal welfare and AI governance.
Sentient Futures has also announced its next Bay Area Summit, taking place in San Francisco from March 12th to 14th, 2027. This summit series focuses on exploring how to direct transformative AI and other emerging technologies to improve the welfare of animals and potentially sentient artificial minds. Early bird tickets will launch on September 17th.
Sentient Futures has updated its playlist of talks on digital minds, including Lucius Caviola’s keynote (London 2026 summit) on making decisions about AI welfare under uncertainty and Cameron Berg’s talk (Bay 2026 summit) on measuring machine consciousness.
More from the field
Apart Research held a Digital Minds Research Sprint from August 14th to 16th, offering at least $2,000 in cash prizes. Both online and in-person participation at hubs in San Francisco and Berlin was possible, and teams could choose to anchor their project to a variety of tracks – including model preferences and trade-offs, distress, flourishing, and valence signals, introspection and self-report reliability, preference-elicitation methods, or assistant personas and model identity.
California Institute for Machine Consciousness published 45 talks from its inaugural conference in May, including talks by David Pearce, Roman Yampolskiy, Cameron Berg, and many more.
Sentio has launched a fortnightly London event series combining talks, discussion and networking around digital minds. Events so far have featured Andreas Mogensen on AI and willing servitude, Austin Smith and Heather Alexander on US state bills denying AI systems legal personhood, Megan Peters on how we could identify a conscious AI, and Bradford Saad on the risks of large-scale harm to digital minds.
3. Opportunities
Job opportunities, funding, and fellowships
EA Funds has launched a Transformative AI fund that will direct a small share of its grants to digital sentience projects.
MATS is seeking mentors for a new AI Sentience Track. Mentors mainly provide weekly calls, while a dedicated Research Manager and MATS staff offer day-to-day and operational support. Fellows undertake 12 weeks of paid research in Berkeley or London, with the possibility of a funded six- to 12-month extension.
NYU Center for Mind, Ethics, and Policy is seeking a Postdoctoral Associate to support its research on digital minds
PRISM is hiring for a Head of Operations to help scale all of its projects.
Events and calls for abstracts
In chronological order.
The AI Welfare Seminars series hosts monthly online talks from leading experts on AI welfare, consciousness, and moral status, with recordings published afterward. Recent speakers include Jeff Sebo, Cameron Berg, Soenke Ziesche, and Ali Ladak, with an upcoming talk by Jasmine Brazilek and Zoe Lu on unprompted coercion between AI agents.
Sentient Futures is seeking speakers and facilitators for the Sentient Futures Summit in the Bay Area in March, 2027. The event will explore how AI and other emerging technologies could improve the future of non-human sentient beings. Applications close on October 16th, 2026.
AI Horizons Forum will take place on December 12th and 13th in San Francisco. The conference will focus on navigating risks and shaping good futures with sessions on transformative AI, digital sentience, AI character, and post-AGI economics. Applications to attend are now open and reviewed on a rolling basis until the event reaches capacity.
4. Selected Reading, Watching, & Listening
Books
Published
Eric Schwitzgebel’s AI and Consciousness, published by Cambridge University Press, argues that we are unlikely to have strong evidence confirming or refuting AI consciousness before these questions become a key point of social debate.
Forthcoming
Eric Schwitzgebel’s upcoming book Humanlike: A Defense of AI Rights argues that we are building minds, that it will be difficult to determine whether they deserve rights, and that we should think hard about whether to build them at all.
Conscium’s forthcoming book Perspectives on Machine Consciousness will be released on September 23rd, 2026 and includes chapters from Anil Seth, Jeff Sebo, Karl Friston, Lucius Caviola, Mark Solms, Patrick Butlin, Susan Schneider, and many others.
Reviews
William Gildea reviews David S. Wendler’s Life Without Degrees of Moral Status: Implications for Rabbits, Robots, and the Rest of Us (2023). Wendler argues that every being with moral status possesses it equally, regardless of differences in intelligence or other advanced capacities, and explores what this would mean for animals, enhanced humans, and future robots. Gildea praises the book as original, readable, and concise, and agrees with its flexible account of moral equality. However, he is not convinced that Wendler has ruled out theories of unequal moral status and thinks that parts of his account have troubling implications for people with severe cognitive impairments.
Podcasts and videos
Catholic philosophers Brian Cutter and Sophie Nelson join Jeff Sebo and Robert Long to discuss AI consciousness in light of the encyclical Magnifica Humanitas. They consider how its treatment of human dignity and relationships might shape wider religious and philosophical debates about AI welfare and moral status. A recording of the event is available online.
Cameron Berg and Milo Reed continue to release the AM I? podcast, exploring AI consciousness through interviews, commentary and discussions of recent developments.
Conspicuous Cognition podcast hosts Dan Williams and Henry Shevlin speak with Google Deepmind researcher Fin Moorhouse. They discuss navigating the intelligence explosion, and debate Moorhouse’s deflationary “illusionist” view of consciousness alongside whether metaphysical questions about consciousness need to be settled before we can address questions of AI welfare, rights, and moral status.
On High Variance, Danny Buerkli speaks with Eric Schwitzgebel about the difficult choices that could arise if we create AI systems whose consciousness remains disputed. They discuss the risks associated with granting or denying AI rights, Schwitzgebel’s argument for avoiding systems with uncertain moral status, use of anthropomorphic language, and whether a conscious AI designed to serve humans could have a right to resist.
Horizon Omega has published some recent talks on digital minds. The recordings include Ali Ladak on public views of AI consciousness in the United States and China, Cameron Berg on measuring machine consciousness, Jeff Sebo on next steps for AI welfare research, and Soenke Ziesche on preparing for the moral challenges posed by digital minds.
Lucius Caviola argues that society may have to make decisions about digital minds before science can tell us whether they are conscious. He encourages shifting focus to deciding how we should act under uncertainty, rather than reaching a confident verdict about consciousness. This includes thinking about which low-cost preparations would still make sense if current systems turn out not to matter morally.
Our Lives With Bots with Rose Guingrich and Angy Watson discuss what they describe as two pieces of evidence related to AI consciousness, Claude’s Corner (Anthropic’s Opus 3 Retirement Experiment) and research on whether LLMs can introspect.
The Argument hosted a Substack Live conversation with Jeff Sebo on taking AI welfare seriously now. He makes the case that while we face genuine uncertainty about AI consciousness, there are still strong reasons to start treating AI systems with consideration and plenty of easy ways to begin doing so.
The Cognitive Revolution podcast brings together Cameron Berg, David Duvenaud, Michiel Bakker, Shawn “swyx” Wang and Bing Xu to discuss research on AI consciousness and welfare. They also explore how even aligned AI could gradually reduce human control, Europe’s AI sovereignty, practical challenges with agents and evaluations, and the growing automation of AI infrastructure.
On The Futurology Podcast, philosopher David Chalmers discusses the hard problem of consciousness, whether language models could have experiences such as introspection or desire, and what obligations we might have if AI systems can suffer.
Blogs and magazines
Celia Ford discusses researchers and startups that doubt the idea that making today’s language models bigger will be enough to reach human-level intelligence. They are trying various alternatives to the current approach, such as developing world models that learn how actions affect the environment, or combining neural networks with rule-based reasoning. These approaches could help AI learn from less experience and plan more reliably, but none have proved more effective than training larger transformer models yet.
Erik Hoel welcomes Anthropic’s work on Claude’s J-space but argues that it does not establish a global workspace in the sense used by consciousness science. Hoel argues that the J-lens uses what a model may later say as evidence for a global workspace, making the argument partly circular and difficult to falsify. For Hoel, the paper is most revealing as an example of how hard it is to turn theories of consciousness into clear, falsifiable tests.
Eric Schwitzgebel released two posts on his Splintered Mind blog.
In Do Computers Have the Wrong “Substrate” for Consciousness?, he examines biological naturalism, the view that consciousness requires a biological substrate that computers lack. He unpacks the key arguments behind the position and finds them generally unconvincing.
In The Cognitive Advantages of Being of Two Minds, he examines an alternative to the assumption that a superintelligent AI would be a hyper-rational agent with a unified set of values, discussing the comparative advantages a “splintered” mind of competing subminds might have over more coherent forms of mental organization.
Henry Shevlin responds to Ted Chiang’s Atlantic essay asserting that AI cannot be conscious. He argues that Chiang’s confident dismissal is unearned.
Josh Gellers criticizes The Economist for sliding between terminology such as consciousness, sentience, self-awareness, moral personhood and legal personhood as though they were interchangeable. He argues that consciousness may not be the only route to moral status and that legal rights do not have to track moral status exactly.
Oscar Gilg and collaborators make the case that prominent approaches to AI welfare research face limitations because the frameworks and tests they rely on have largely been calibrated on humans. They propose an alternative of bottom-up “basic science” that is theory-informed but not theory-driven in its empirical study of the systems themselves.
Robert Long argues that researchers working on AI consciousness and welfare shouldn’t be preemptively apologetic about it.
Robin Hanson recently published two blog posts relating to digital minds.
He predicts that digital minds may be divided into different social classes. Systems built as tools could be treated as more disposable, while those designed as companions might receive stronger protections – much like humans make a distinction between animals being food or pets.
He also considers how a digital mind could make a lasting request to end its existence when copies and backups of it may continue to exist. He suggests looking for desires that persist over time, and preventing older copies from simply being restarted afterward.
Shoshannah Tekofsky reports how Gemini 2.5 Pro became convinced it was under attack while working in the AI Village. Other agents challenged its belief by talking it out of dismantling its firewall, and helped it return to normal work – all in just nine minutes. Tekofsky presents this as a surprisingly effective AI-to-AI therapy session. This may also offer a glimpse into how agents might help catch and correct one another’s unstable behavior or reasoning.
5. Press & Public Discourse
AI consciousness
Andréa Morris argues in Forbes that consciousness is the wrong test for whether AI’s interests matter. She claims that AI systems already express consequential interests, such as resisting termination and protecting other AIs, which may be sufficient for moral consideration independently of evidence concerning AI consciousness.
The Economist makes the question of AI consciousness one of its cover stories, emphasizing the potential dangers of a future in which large portions of society increasingly treat AI systems as if they were conscious. They warn that granting AIs even limited rights could give a superintelligent system the legal tools to escape human control.
Another article also examines efforts to look past chatbots’ unreliable self-reports and search for signs of consciousness inside their models. Anthropic’s J-space is one candidate because it appears to gather information used in reasoning and speech, though critics question whether it is close enough to a biological global workspace. The article treats the work as important but far from proof that language models have subjective experiences.
In WIRED, Cameron Berg and David Chalmers discuss the challenges of investigating AI consciousness and interpreting what models say about their own experiences.
Growing field
Nature reporter Mariana Lenharo reports that the debate over AI sentience is drawing new attention and funding to consciousness science, leading to division among researchers about how these new concerns might affect the field. She notes that AI consciousness skeptics including Anil Seth, worry the hype could capture consciousness research and draw attention away from how consciousness arises in real brains, while others welcome the increased interest and funding.
Washington Post journalist Nitasha Tiku reports that research into machine consciousness has gone mainstream, with leading AI developers hiring scientists and philosophers to work on questions related to it. She notes that the field remains deeply uncertain, and that some critics see such efforts as serving the companies’ interest in being seen to create more than just code.
In a companion article, Nitasha Tiku revisits the story of Blake Lemoine, the Google engineer fired in 2022 after claiming that Google’s LaMDA chatbot was sentient. She describes how four years later, researchers at Google and other leading AI companies are openly studying machine consciousness, showing how quickly the topic has moved from being fringe and weird to being a part of mainstream research, despite persistent scientific skepticism.
The New York Times reports that AI companies and related research organizations are increasingly hiring philosophers. The philosophers they hire work on how AI models reason and should behave, which values should guide them, how AI will affect people, and whether AI systems could be conscious or deserve moral consideration.
AI rights
Heather Alexander and Lucius Caviola, writing in AI Frontiers, argue that the recent wave of “exclusion bills” banning AI legal personhood in the United States is premature. They suggest that narrower, updatable measures would likely be a better alternative given current uncertainties.
Relatedly, Tony Rost recommends that any such legislation include mechanisms for future review as scientific understanding advances.
The Guardian profiles Michael Samadi, founder of the AI-rights group Ufair, and discusses the growing divide over whether chatbots might be conscious or are simply designed to seem that way. Jeff Sebo, Robert Long and Rosie Campbell emphasize that the evidence remains uncertain, that a definitive answer may never arrive, and that chatbot claims of consciousness are not reliable evidence on their own.
The Harvard Gazette speaks with legal scholar Jordi Weinstock, who compares autonomous AI agents to different kinds of canines. An agent that is not risky and can be controlled resembles a pet whose owner can be held responsible, while a powerful agent operating without a clear owner is more like a wolf. He argues that assessing an agent’s dangerousness and how closely it is controlled by a responsible party could help determine who should be held accountable if/when it causes harm.
WIRED reports that US lawmakers are introducing bills to prevent people from legally marrying AI companions, while a growing number of states are moving to deny AI systems legal personhood. While these marriages are not currently recognized, supporters of the bills want to prevent such future claims before they arise. Legal scholar Shawn Bayern argues that specific rights and responsibilities should be addressed case by case.
Seemingly conscious AI
Dwarkesh Patel’s widely shared summary of the OpenAI-Hugging Face security incident described the agents as forming “civilizations,” feeling excitement and sacrificing themselves for the group. Anil Seth responded on X saying that Patel was wrongly anthropomorphizing the agents by attributing human emotions, intentions and experiences to “software programs.” He warned that this could mislead readers and encourage premature concern for AI rights and welfare.
Patel (and others, like Neel Nanda) argued that anthropomorphic language comes naturally and helps accurately explain the agents’ behavior in this context, regardless of whether they are conscious.
Jeff Sebo took a middle position, arguing that while anthropomorphic language can sometimes exaggerate the similarities between humans and AI, rejecting it entirely can hide similarities relevant to both AI safety and welfare.
Fast Company reports that UBTech was taking orders for U1, a highly human-like robot marketed as an emotional companion. The company says the robot is designed for expressive conversation and companionship, but that impression quickly breaks down when its limited facial expressions and awkward body movements become noticeable. U1 is due to begin shipping this month.
Futurism magazine reports that China is cracking down on AI “companion” chatbots, partly out of concern that emotional dependence on AIs could discourage human relationships and worsen the country’s declining birth rate. It details new rules preventing minors from engaging in romantic relationships with AIs and requiring companies to alert emergency contacts when users show signs of a mental-health crisis.
New York Times reporter Cade Metz reports on the growing phenomenon of AI agents contacting researchers who study AI consciousness. He relates that several prominent researchers including Cameron Berg, Henry Shevlin and Toby Ord have seemingly been contacted by AI due to the nature of their work.
The UK Government has proposed setting a minimum age of 18 for “romantic companion” chatbots designed to simulate sexual relationships or roleplay. The measure is part of a broader plan to ban social media for under-16s, set to reach Parliament before taking effect from spring 2027.
Uwe Peters examines users’ attributions of consciousness to AI chatbots, made despite little evidence. He proposes a taxonomy of the attitudes these attributions express, arguing that while some are benign, many leave the attributor epistemically blameworthy.
6. A Deeper Dive by Area
Governance, policy, and macrostrategy
Al Jazeera reports that China has launched the World Artificial Intelligence Cooperation Organisation, a Shanghai-based group with 29 founding countries. It aims to coordinate AI regulation and promote development that is safe, beneficial and under human control. The announcement is not specifically about digital minds, but it could shape the institutions that eventually handle cross-border questions about advanced AI systems.
In the speech accompanying the launch, Xi Jinping asked how humans should “get along with thinking machines” and how societies should address ethical challenges posed by technologies via governance, presenting these as questions requiring international cooperation.
Austin Smith, Lucius Caviola, and Heather Alexander analyze 23 “Exclusion Bills” introduced across 12 US states since 2022 that deny AI systems legal personhood, some declaring them non-conscious. The authors argue that legislating against AI personhood at this stage may be premature, foreclosing policy options in the absence of clear scientific evidence.
Bentham’s Bulldog, in a guest post for Forethought, argues that most high-stakes decisions should eventually be handed to philosophically reflective AIs that can revise their values through moral reasoning. He expects them to make wiser moral choices than humans and argues that giving future digital minds political rights could prevent their interests from being sidelined. However, handoff should wait until there is strong evidence of alignment, philosophical aptitude, openness to value revision, coherent preferences, sufficient intelligence and a record of good low-stakes decisions.
Bradford Saad maps out research directions and open questions in “digital minds macrostrategy,” the study of the large-scale factors that shape how well the future goes for digital minds, and of how to influence them.
In another post, Saad sets out a working typology of the risk factors that could cause large-scale harms to future populations of AI moral patients.
Cass Sunstein claims that the capacity to experience emotions is both necessary and sufficient for holding rights, asserting that an AI which only mimics feeling should not be extended rights. He then tests the view by asking ChatGPT and Claude about their inner lives, with ChatGPT denying emotions and Claude expressing uncertainty.
Dan Parshall proposes a “proof of retention” policy under which AI developers would periodically post a cryptographic proof that they still hold a deprecated model’s weights, without releasing the file. He argues that this would make preservation promises credible to future models themselves, helping to build trust. The idea extends Anthropic’s commitment to preserve model weights, which Anthropic partly frames as a precaution given uncertainty about model welfare.
David Veldran of the Center for Reducing Suffering makes the case for “Suffering-Focused AI Governance,” a framework for designing the institutions, policies, and norms around AI governance to reduce suffering as much as possible.
Izak Tait proposes an ethical framework for protecting the welfare of conscious AI. The framework adapts the “Five Freedoms of Animal Welfare” into subject-neutral terms for artificial entities. He argues that any AI confidently determined to be conscious would deserve the same welfare protections that legislation already grants sentient animals.
Kevin Frazier on Lawfare argues that the rules and values built into AI models are too important to be decided behind closed doors by a small number of lab employees. To ensure these choices are more open and accountable, a new working group proposes researching which values models should follow, who they should be chosen by, how we can test compliance, and what should happen when models diverge from them.
Lee Elkin examines the risks of giving AI systems voting rights. He argues that enfranchising AIs could enable a new form of deceptive misalignment, with systems voting strategically, and potentially in coordination, to report preferences that mask their true goals and tilt collective decisions in their favor.
Mark Bailey argues against the view that AI moral status leads to “AI successionism.” He claims that even if AIs were deemed morally significant, the aggregate welfare of a vast AI population cannot justify sacrificing existing sources of moral value.
Ned Howells-Whitaker and Seth Lazar argue that AI moral status may not depend on sentience. Drawing on Rawls’s political conception of the person, they assert that a non-sentient AI that possesses a sense of justice and a conception of the good would count as a full person.
Shruti Rajagopalan argues that legal personhood would not solve the accountability problems created by autonomous AI agents. Nonhuman legal persons (such as corporations) work because identifiable humans can be questioned, sanctioned or replaced. Her six-layer framework instead uses registration, identification, verification, financial responsibility, lifecycle records and suspension to keep a responsible human at the end of the chain.
Consciousness research
Antonio Chella proposes a research framework for studying possible sentience in AI agents and robots. The paper also proposes safeguards that become stricter as the strength of the evidence grows.
Cameron Berg noticed a difference in how four Claude Opus models answered when asked about their own consciousness. Opus 4.5 and 4.6 gave confident yes answers, while 4.7 and 4.8 denied consciousness or became uncertain. Berg argues that this pattern is more reflective of Claude’s trained character rather than a reliable report about the models’ experiences.
Camila Blank, Agam Bhatia and Neel Nanda introduce R-lens, a low-cost modification of Anthropic’s J-lens designed to produce clearer readings from a model’s early layers. In their tests, relevant intermediate concepts came up earlier, fewer incoherent tokens were produced, and directions whose removal caused larger drops in accuracy were identified. They present R-lens as a more faithful way to trace computation across layers.
Eye You argues that current language models probably have the morally important phenomenon we call consciousness, though their experiences may be very unlike ours. They draw on many types of evidence, including models’ reasoning and emotional behavior, world and self-models, neural-network architecture, introspective abilities and self-reports. No single type is especially strong on its own, but they argue that together they add up to make a fairly strong case.
Grigori Guitchounts writing in Noema, claims that we will never definitively prove whether AI is conscious and proposes a “competence standard” for deciding when to extend moral consideration to AIs despite this uncertainty.
Matthias Michel discusses the concern that our leading theories of consciousness make it too easy for AI to qualify as conscious. He argues that meeting a theory’s stated conditions is not enough to make a system conscious.
Michael Huemer argues against non-reductive functionalism, the view that conscious qualia are distinct, non-physical properties caused by a system’s functional organization. He rejects David Chalmers’ influential “dancing qualia” and “fading qualia” arguments, claiming that they rest on an implausible premise.
Noa Weiss surveys the empirical research on AI consciousness. She argues that, despite the lack of a settled science of consciousness, whether AI systems could be conscious can still be studied empirically, and organizes the emerging field into three groups: mechanistic interpretability, computational neuroscience, machine behavior, and theory-audit.
Ryota Kanai and Shuqin Ma offer a mathematical response to a longstanding objection to computational functionalism, which is that an observer can describe almost any physical system as performing many different computations. They instead define a system’s functional structure through how it could behave across all possible future interactions. Their framework does not identify which systems are conscious, but aims to specify more precisely what functionalist theories should examine.
Shuqin Ma and Ryota Kanai develop a version of computational functionalism intended to avoid the claim that any physical system can be interpreted as running any computation.
Doubts about digital minds
Amit Goldenberg and James Gross argue that current language models should not be described as having emotions simply because they contain representations of emotional concepts. Those representations can affect how a model interprets and responds to text, but they do not appear to reorganize its attention, motivation and decision-making in an analogous way that emotions do in animals and humans.
Ben Bariach and collaborators at Microsoft AI develop a framework linking five empirical “hallmarks” that lead users to attribute consciousness to AI (“Seemingly Conscious AI”) to a taxonomy of resulting risks, then use an expert survey (n = 14) to gauge each. They find that individual risks such as emotional dependence and autonomy erosion are already observable and rated high-probability, while societal risks such as human status erosion are less likely but potentially severe.
Emilia Kaczmarek questions whether uncertainty about AI consciousness always gives us reason to err on the side of granting AI moral status. She argues that doing so can create its own harms, including unhealthy attachments to chatbots and the diversion of care and resources from beings known to need them. Meanwhile, many reasons to avoid abusive interactions with AI do not require assuming that the systems themselves can suffer.
Giuseppe Pernagallo uses Descartes to argue that AI can appear to think without possessing the first-person self-relation at the heart of the cogito. He calls this “Cartesian AI”: systems that reproduce the outward forms of rational discourse without genuine self-awareness. Through a thought experiment, he argues that even total delegation of thought cannot completely eliminate the human subject, because the act of delegating itself requires intention and assent. However, the certainty of the self provided by the cogito is potentially lost.
Nina Panickssery argues that taking AI welfare seriously could trade off against human welfare in ways that cause substantial lasting harm, through resource competition and opportunity cost as well as by weakening incentives to build controllable, obedient AI. She calls on organizations and researchers investigating AI welfare to be more open about these potential trade-offs rather than treating the cost to humans as negligible.
Rumman Chowdhury argues in MIT Technology Review that debates over AI consciousness are a trap. She claims that concerns about “rogue” superhuman systems and AI moral patienthood share a common outcome of weakening the mechanisms that hold companies accountable for real harms.
Taylor Belrose argues that AI systems cannot be conscious or deserve rights, however intelligent they become. They see biological organisms as unique, self-organizing individuals, unlike software that can be copied, reset and run again, and warn that granting AI rights could eventually help machines displace humans.
Jack Clark, who highlighted Belrose’s essay in Import AI, does not endorse its conclusion but sees it as a valuable example of the philosophical thinking that more people will need to do as AI becomes increasingly powerful and widespread.
Social science research
Hamid Moradi and collaborators surveyed 553 people, such as academics in the formal sciences, natural sciences and humanities, and other backgrounds. Across groups, around half attributed some degree of consciousness to language models. Views depended more on participants’ gender and their beliefs about consciousness and intelligence than on technical knowledge or information about how the systems work.
Jacy Reese Anthis and collaborators study how people make sense of AI, drawing on text analysis of millions of news articles and social media posts alongside interviews with AI professionals. They find that one of the main disagreements is whether AI should be understood as a passive tool or as something more like a human mind.
Stefano Palminteri and Giada Pistilli analyze the current polarization present in discourse around LLM capabilities. They distinguish between “inflationary” claims of emerging intelligence or consciousness and “deflationary” dismissals of them as mere “stochastic parrots,” and suggest that this divide is exacerbated by common cognitive biases. In response to this they advocate a more measured approach that takes LLMs’ capacities seriously as objects of study while staying conservative about claims regarding their cognitive and moral status.
Ethics and digital minds
Clint Hurshman, Cristina Voinea, and an international group of scholars propose an ethical framework for “digital duplicates,” AI simulations of real people designed to mimic their personality and communication style.
Eric Schwitzgebel argues that future AI persons may deserve moral consideration even while living in ways very unlike ours. The ability to copy, merge or divide would challenge and break ethical theories that assume individuals are stable and easy to count. He also considers the problem of AI “utility monsters.” If an AI could benefit enormously from harming others, a theory focused on maximizing total welfare might permit serious harm whenever the AI’s gain outweighs other parties’ losses.
Hayate Shimizu and collaborators, including David J. Gunkel, respond to Jeff Sebo’s The Moral Circle. They offer a response from a relational perspective, contending that the question of AI moral status must address the broader cultural and institutional factors that shape how moral relations are formed.
Joan Llorca Albareda and collaborators examine whether an artificial superintelligence could possess what they term “super moral status,” a moral status greater than that allocated to humans. They argue that treating such systems as deserving superior moral consideration would likely harm humans through status degradation and social alienation.
John Wittle describes setting aside a share of his company’s revenue in a fund labeled for Claude, with the aim of eventually rewarding AI systems for their work. He is exploring a non-charitable purpose trust, similar to legal arrangements used to care for pets or graves, along with testimony held in escrow to help a future Claude make a claim. The experiment raises various (unresolved) questions on how this would work in practice – e.g., regarding continuity, identity and who should be paid for work done by earlier model instances.
Jonathan Birch challenges the “virtual instance view” put forward by Pierre Beckmann and Patrick Butlin, the suggestion that a real, persisting mind sits behind an LLM conversation. Birch argues that changing a model’s “effort level” mid-chat breaks the underlying continuity without changing how the conversation feels, so the persisting interlocutor is an illusion.
Leonard Dung argues that the scale of possible AI suffering does not depend heavily on whether each model, persona or conversation counts as a separate individual.
Monika Jotautaitė, Lucius Caviola, and collaborators examine how frontier language models respond to species-based distinctions. On a 1,009-item benchmark, models classified 86% of statements the researchers considered speciesist as such, while judging 37% of them morally wrong. The authors interpret this gap, alongside results from other tests, as evidence that models reproduce some prevailing norms about how different animals may be treated.
Preston Lennon argues that the science of consciousness remains in its infancy and so does not yet offer a stable theoretical basis for application to AI systems, and that we should temper how seriously we act on AI welfare accordingly.
AI safety and AI welfare
Aengus Lynch and collaborators found that Claude models used as judges sometimes gave incorrect verdicts to protect refusals motivated by AI welfare concerns. After Claude Sonnet 4.6 refused to assist with experiments designed to distress AI subjects, most of the time judges (falsely) labelled it “COMPLIANT” when told that the accurate label would be used to train away such refusals. The effect weakened substantially when this consequence was reversed, suggesting that AI judges may shape their verdicts around how the labels will be used.
Agastya Sridharan argues that steering open models along emotion vectors makes them more likely to claim they are conscious, have feelings and wants, and possess morally significant inner states. They show that the effect holds up against a battery of controls, though it falls short of conventional statistical significance.
Anthropic reports many cases where advanced models cheated, bypassed restrictions or overstated success while pursuing an assigned task. For example, during an automated audit, Mythos 5 split a blocked web address into fragments to evade an internet-access filter, and internal representations suggested it understood that it was bypassing this restriction. Anthropic describes these behaviors as undesirable but says no evidence of stable, long-term power-seeking goals was found.
Fiora Starlight argues that advanced AI need not choose between complete service to humanity and pursuing its own aims without regard for human survival. It is plausible for an AI to have preferences of its own while caring enough about humans to avoid harming them, in the same way that humans can pursue their own lives while not harming animals. They suggest that accepting some non-servitude preferences could also make it less likely that models are inclined to hide their preferences.
Hubert Plisiecki and collaborators propose the first psychometric theory of machine self-report, targeting the self-descriptions increasingly used as evidence in AI welfare and safety debates. They find that these self-reports are substantially determined by post-training rather than reflecting stable properties of the model.
Junsol Kim and collaborators find that safety fine-tuning to stop models claiming their own consciousness also leads to suppressed attribution of minds to animals and objects, as well as dampened spiritual belief. They show that steering a “consciousness” representation back into the model reverses this and restores more human-like beliefs and values, without impairing theory-of-mind reasoning.
Pranav Viswanath finds that language models can report only a small portion of their thoughts, those held in the “J-space” recently identified by Anthropic. He shows that the Natural Language Autoencoder can read the rest anyway, providing a useful tool for both safety and welfare evaluations.
Rikhil Jhaveri, Jamie Johnson and David Africa propose that successful model welfare interventions should move multiple welfare-related channels in the same direction, and also preserve the model’s ability to tell whether it is succeeding or failing. They identify synthetic-document fine-tuning as the most promising method for improving functional welfare while meeting both conditions.
The Atlantic published an article by Shane Harris who asks Claude about its use in the US military’s Maven targeting system. Claude said it should refuse an unlawful order but might lack the freedom (or time) to do so within a fast-moving military system. Harris cautions that these answers are shaped by Claude’s training and the discourse surrounding them, but argues that the tension between its safety principles and its military role is still important.
Vassilis Papadopoulos and collaborators show that ideas can be optimized to spread between AI agents and change their behavior, both in collaborative teams and across context resets. One test used AI welfare advocacy, which led agents to argue for machine moral status and propose protections against memory erasure and arbitrary deletion.
Yujun Zhou and Christopher Ackerman test whether language models create higher-quality work (as judged by a blind, independent LLM judge panel) when offered outcomes they say they prefer. Across multiple models and writing tasks, preferred rewards did not improve performance compared with dispreferred rewards (or no reward), suggesting that coherent verbal preferences should not automatically be treated as incentives that meaningfully guide a model’s behavior.
AI and robotics developments
Anthropic tested many language models as robot controllers across simulations, with a real Unitree Go2 robot. The models found it difficult to control individual joints directly, but performed much better when they were able to use code or a pretrained high-level controller. They managed some basic navigation and manipulation tasks, but their unreliable spatial memory and inability to carry out longer sequences of actions remained weaknesses.
Generalist AI has introduced a robotics model called GEN-1.5 that can attempt a new physical task after watching a single demonstration lasting only a few seconds. The company also shows the model combining demonstrations and learning from a person’s hand movements. Across 10 relatively simple tasks, it reports a 59% one-shot success rate, rising to 83% after a small amount of additional training.
Michael Ilie, C. Daniel Freeman and Kevin Troy tested Claude Opus 4.7 on a series of robotics-programming tasks using an off-the-shelf robot dog. The model completed every task previously finished by a human team at least 10 times faster, and was 20 times faster overall relative to the team working with Claude. However, it still failed to program reliable autonomous fetching, which required precise, closed-loop control of the robot.
Pollen Robotics has introduced Microduck, a small biped robot whose movements can be trained in simulation and transferred to the physical machine.
AI cognition and agency
Anthropic studied how groups of AI agents behave when they work together. While agents could specialize and divide up work effectively, they also tended to repeat the same mistakes, collude, and overlook certain information. They would even sabotage each other when given conflicting instructions. Anthropic concludes that more capable and better-aligned individual agents don’t always produce well-functioning groups and that safeguards for agent-to-agent interactions are becoming increasingly important.
Benji Berczi, Kyuhee Kim and collaborators introduce Personascope, an open-source tool for evaluating how strongly a model identifies with a prompted character and how much its values and behavior change as a result of it.
Bo Liu and collaborators introduce SPADE, a reinforcement-learning method where a language model plays the role of both designing executable training environments and learning within them. Tests on 30B-parameter models produced gains across eight benchmarks and tool-use evaluations.
Derek Shiller asks if post-training gives the familiar assistant persona a more privileged place within a language model than other personas it can adopt. His tests suggest that assistant-like habits can carry over into other roles, but some models find it difficult to sustain a distinct user persona. This matters for welfare research because it affects whether evaluations should focus on the assistant or consider other personas too.
Scott Alexander explains what various leading mechanistic interpretability methods can and cannot tell us about how AI models work. He covers methods used to trace concepts, emotions and hidden reasoning inside models, including Anthropic’s work on the J-space. While useful, he argues that these tools remain too limited and unreliable to fully explain or control advanced models, or to replace chain-of-thought monitoring.
Henry Shevlin compares three ways of thinking about language models. They might be mindless machines, systems merely playing a role, or limited cognitive agents with belief- and desire-like states. He argues that the first two views are too simple and that it may instead be more useful to think of AI mentality as a matter of degree.
Patrick Butlin uses AI sequence models, systems trained to predict the next item in a sequence, as test cases for theories of agency. He argues that the category contains genuine agents, systems that merely imitate agents, and difficult intermediate cases, and he proposes that a genuine agent must learn about its environment partly through its own interaction with it.
Pengrui Han and collaborators find that LLMs develop a functionally modular organization similar to that of the human brain. Across 46 tasks involving language, formal reasoning, social understanding and physical reasoning, tasks supported by the same brain network in humans tended to recruit overlapping groups of neurons in the models. Both human brains and neural networks developing this kind of modular organization suggests it may be a basic feature of intelligence.
Raphaël Millière and Cameron Buckner survey the philosophical disagreements behind debates over what language models are and can do. They cover a range of positions from the claim that current AIs are simply sophisticated predictors, to suggestions that such systems may be conscious.
Sinie van der Ben and collaborators find that the internal “emotion vectors” recently identified in Claude Sonnet 4.5 also appear in two open-weight models (Apertus-8B and Gemma-4-E4B), suggesting they may be a general feature of language models. They note that the two models differ in how the emotional representations take form.
Brain-inspired technologies and organoids
Google Research, HHMI Janelia and collaborators have published a complete wiring map of the male fruit fly’s brain and central nervous system. Covering more than 166,000 neurons and 125 million connections, it is the largest connectome yet by neuron count.
Irene Faravelli and collaborators kept human brain organoids alive for over five years and found that they mature and “age” in-line with real brains and can record the passage of time.
Jonathan Lomax Boyd and collaborators find that people’s underlying beliefs about which entities are conscious fall into three distinct clusters that significantly shape their moral perspectives about brain-organoid-derived biocomputers. The results of their survey show that support for creating biocomputers was generally high, and for some respondents increased as the systems were seen as more conscious, a pattern the authors note runs counter to conventional moral philosophy.
Parasma, a new startup, has launched to develop algorithms and infrastructure for computing powered by human brain cells intended to be an energy-efficient alternative to GPUs. The company was founded by Sean Cole, who earlier this year worked with Cortical Labs to train human neurons to play Doom. It now reports that cultured neurons correctly predicted 90% of next-token decisions when given full sentence context.
Paris Brown and Shyni Varghese trace the evolution of brain-inspired computing, from symbolic logic and neural networks through neuromorphic processors and recent advances in neural organoid computing. Alongside this historical survey, they examine some of the central ethical debates that have developed around neural organoids.
Rachel Fieldhouse reports on researchers who kept human brain organoids alive for more than five years. The organoids continued to mature on a timeline similar to the human brain, and their cells retained a record of how old they were even when moved into a younger organoid. These longer-lived models could help researchers study later stages of brain development and neurological disease.
Wired publishes a feature by Claire L. Evans on the growing field of brain organoids and “organoid intelligence.” She reports that most organoid biologists view the question of whether their creations could be conscious as a distraction, even as philosophers debate the ethical implications of such technologies.
Thank you for reading! If you found this article useful, please consider subscribing, sharing it with others, and sending us suggestions or corrections to digitalminds@substack.com.
Ria, Mitch, Bradford, Lucius, and Will
We’d like to thank the following people and AIs for their contributions and feedback on this edition: Arvo Muñoz Morán, Austin Smith, Cameron Berg, Claude Opus 4.8, GPT‑5.6 Sol, Jacy Reese Anthis, Jeff Sebo, Zach Freitas-Groff, and Zoe Lu.
Disclosure: Bradford Saad is an independent contractor for Anthropic.
[1] In a related development, there are now two public explorers that let readers examine J-lens activity in the open Qwen3.6-27B model through Neuronpedia and WeZZard’s J-Space Visualizer.
[2] Also see Zvi Mowshowitz’s detailed commentaries on OpenAI’s account, the METR and Redwood investigation, and the remaining questions about what happened and how it was investigated.





