clic lab logo clic lab

Selected Publications

* denotes equal contribution.

Large language models persuade without planning theory of mind

Moore, J., Overmark, R., Cooper, N., Cibralic, B., Haber, N., & Jones, C. R.

Cognitive Science (accepted) 2026

LLMs are effective at persuading people despite being poor at representing and reasoning about their partner's beliefs, suggesting that their persuasion does not rely on planning theory of mind.

Language statistics and false belief reasoning: Evidence from 41 open-weight LMs

Trott, S., Taylor, S. M., Jones, C. R., Michaelov, J. A., & Rivière, P. D.

Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics 2026

Compares 41 open-weight language models on false belief tasks, asking how far performance can be predicted from properties of the training data rather than from any particular model.

LLMs and people both learn to form conventions — just not with each other

Jones, C. R., Lombardi, A., Mahowald, K., & Bergen, B. K.

Proceedings of the Annual Meeting of the Cognitive Science Society, 48 2026

People and LLMs each converge on shared conventions over repeated reference games, but pairs made up of one human and one model do not.

Developmental trajectories of situation modeling and mentalizing in transformer language models

Rivière, P. D., Jones, C. R., & Trott, S.

preprint 2026

Follows when situation models and mental state inference appear over the course of training, rather than testing only the finished model.

Inverse Turing Bench: Evaluating language models as judges of human vs. AI dialogue

Hager, W., Rathi, I., Hasan, M., & Jones, C. R.

preprint 2026

Turing test work usually asks whether a model can pass as human. We ask the reverse question: how well language models can tell human and AI dialogue apart.

Implicit vs. explicit prompting strategies for LVLMs in referential communication

Zeng, P., Paige, A. J., Li, W., Brennan, S. E., Rambow, O., & Jones, C. R.

preprint 2026

When people refer to the same thing repeatedly, they settle on shorter descriptions without being told to. We compare prompting strategies that make that process explicit with ones that leave it implicit.

Large language models pass a standard three-party Turing test

Jones, C. R., & Bergen, B. K.

Proceedings of the National Academy of Sciences 2026

In a pre-registered three-party Turing test, GPT-4.5 prompted to adopt a human-like persona was judged to be the human 73% of the time, significantly more often than the real human participants.

AI epistemic risks: Emerging mechanisms and evidence

Yang, M., Casper, S., Stray, J., Li, J., Jones, C. R., Gausen, A., … & Pelrine, K.

preprint 2026

A review of the mechanisms by which AI systems might change how people form and revise beliefs, and the evidence available for each.

Lies, damned lies, and language statistics: A comprehensive review of risks from manipulation, persuasion, and deception with large language models

Jones, C. R., & Bergen, B. K.

Artificial Intelligence Review 2026

Reviews the empirical work on how far LLMs can persuade and deceive, sets out the risks that follow, and considers what might be done about them.

How open must language models be to enable reliable scientific inference?

Michaelov, J. A., Arnett, C., Chang, T. A., Rivière, P. D., Taylor, S. M., Jones, C. R., … & Altman, M.

preprint 2026

Research that treats language models as objects of study depends on knowing what went into them. Sets out which aspects of openness, from weights through to training data, matter for drawing reliable conclusions.

International AI Safety Report 2026

Bengio, Y., Clare, S., Prunkl, C., … Jones, C. R., et al.

Expert report 2026

The second full edition of the international expert report on the capabilities and risks of advanced AI.

International AI Safety Report 2025: First key update — capabilities and risk implications

Bengio, Y., Clare, S., Prunkl, C., … Jones, C. R., et al.

Expert report 2025

Updates the evidence on AI capabilities and the risks that follow, including recent gains from reasoning methods and from extra computation at inference time.

Do large language models have a planning theory of mind? Evidence from MindGames: a multi-step persuasion task

Moore, J., Cooper, N., Overmark, R., Cibralic, B., Haber, N., & Jones, C. R.

Conference on Language Modeling (COLM) 2025

Introduces MindGames, a task in which agents have to work out what their partner believes and wants in order to persuade them. People do better than o1-preview, though the model does well when little mental state inference is needed.

Judging the judges: Displacing and inverting the Turing test to investigate the interrogator

Rathi, I., Bergen, B. K., & Jones, C. R.

Proceedings of the Annual Meeting of the Cognitive Science Society, 47 2025

Turns the Turing test around to ask what makes a good interrogator, rather than a convincing witness.

Dissecting the Ullman variations with a SCALPEL: Why do LLMs fail at trivial alterations to the false belief task?

Pi, Z., Vadaparty, A., Bergen, B. K., & Jones, C. R.

Proceedings of the Annual Meeting of the Cognitive Science Society, 47 2025

Isolates which alterations to the classic false-belief task break LLM performance, and why.

Does language stabilize quantity representations in vision transformers?

Rivière, P. D., Parkinson-Coombs, O., Jones, C. R., & Trott, S.

Proceedings of the Annual Meeting of the Cognitive Science Society, 47 2025

Asks whether vision transformers trained alongside language hold more stable representations of quantity than models trained on images alone.

Prompt engineering large language models' forecasting capabilities

Schoenegger, P., Jones, C. R., Tetlock, P. E., & Mellers, B.

preprint 2025

Tests whether prompting strategies drawn from the forecasting literature improve the accuracy of language model predictions.

People cannot distinguish GPT-4 from a human in a Turing test

Jones, C. R., Rathi, I., Taylor, S., & Bergen, B. K.

ACM Conference on Fairness, Accountability, and Transparency (FAccT) 2025

Interactive, controlled Turing tests in which GPT-4 is judged human at rates comparable to human baselines.

When large language models are more persuasive than incentivized humans, and why

Schoenegger, P., Salvi, F., Liu, J., Nan, X., Debnath, R., Fasolo, B., Jones, C. R., … & Karger, E.

preprint 2025

Compares language models with human persuaders who were paid for succeeding, and looks at what accounts for the difference between them.

GPT-4 is judged more human than humans in displaced and inverted Turing tests

Rathi, I., Taylor, S., Bergen, B. K., & Jones, C. R.

Workshop on Detecting AI-Generated Content (GenAIDetect), COLING 2025

Two non-interactive Turing test variants in which the best GPT-4 witness is judged human more often than actual human witnesses. Best Paper Award.

Comparing humans and large language models on an experimental protocol inventory for theory of mind evaluation (EPITOME)

Jones, C. R., Trott, S., & Bergen, B. K.

Transactions of the Association for Computational Linguistics 2024

Introduces EPITOME, a battery of theory of mind tasks drawn from the experimental literature, and compares human and model performance across them.

Do multimodal large language models and humans ground language similarly?

Jones, C. R., Bergen, B. K., & Trott, S.

Computational Linguistics 2024

Across four pre-registered studies we adapt techniques developed to investigate embodied simulation in humans, asking whether multimodal LLMs are sensitive to sensorimotor features that are implied but not explicit in descriptions of an event.

Does GPT-4 pass the Turing test?

Jones, C. R., & Bergen, B. K.

Proceedings of NAACL 2024

A large set of interactive Turing tests asking whether people can distinguish GPT-4 from a human under controlled conditions.

Multimodal language models show evidence of embodied simulation

Jones, C. R., & Trott, S.

LREC-COLING 2024

We adapt a task designed to test embodied simulation in human comprehenders, and find that multimodal language models are sensitive to some visual properties that a description implies but never states.

Does word knowledge account for the effect of world knowledge on pronoun interpretation?

Jones, C. R., & Bergen, B. K.

Language and Cognition 2024

People draw on knowledge about the world when they decide what a pronoun refers to. We ask how much of that effect can be accounted for by knowledge about words alone.

Do large language models know what humans know?

Trott, S.*, Jones, C. R.*, Michaelov, J. A., Chang, T. A., & Bergen, B. K.

Cognitive Science 2023

Compares LLMs and humans on a false-belief task to ask how far distributional language statistics alone can account for human-like mental-state attribution.