LaTeX editing, your library, and agents that actually read the literature — in one window. This is the real interface: click around.
Large language models increasingly succeed on tasks once thought to require social reasoning, yet their performance on false-belief tasks remains brittle. We hypothesize that much of this brittleness stems from a failure to anchor reasoning to a specific agent's knowledge state, rather than from a lack of underlying capability.
We test this hypothesis with a lightweight intervention that asks the model to restate what each character can and cannot observe before answering. The result is a large, consistent improvement that requires no additional parameters.
Given a scenario and a question about agent a, we prompt the model to enumerate the events a witnessed, mark events hidden from a, and only then resolve the query conditioned on a's partial observation. This decomposition mirrors the sub-steps humans perform implicitly.
End-to-end prompting conflates the beliefs of different characters, causing errors that compound as stories grow longer. We argue that an explicit, symbolic representation of "who knows what" is both interpretable and more robust than relying on the model's latent state.
For each character we instantiate a graph whose nodes are propositions and whose edges encode observation events. Queries are answered by reading off the relevant character's sub-graph, sidestepping the interference that plagues monolithic prompting.
LLMs such as GPT-4 seemingly excel at tracking characters’ beliefs in stories, yet struggle to translate this capability into strategic action. The core challenge lies in identifying the implicit inferences about mental states — without being explicitly asked, as in ToMi — that lead to choosing the correct action.
The foresee step prompts the model to anticipate likely future events and the challenges facing each character; the reflect step reasons over candidate actions against those futures before committing. FaR generalizes to diverse out-of-distribution story structures and scenarios.
We use the Sally–Anne unexpected-transfer task and the Smarties unexpected-contents task, each with carefully constructed controls to rule out shallow lexical shortcuts. Items are presented one sentence at a time to probe incremental belief updates.
Emergence of ToM-like behavior may be an unintended by-product of training on human-generated text. We caution against both over-attribution of mental states and premature dismissal, and call for adversarial benchmarks that separate the two.
Agents skim and read the papers in your library, then edit main.tex. New papers enter only through the arXiv API — so agents can only cite papers added that way, or by you.
Large language models increasingly succeed on tasks once thought to require social reasoning, yet their performance on false-belief tasks remains brittle. We hypothesize that much of this brittleness stems from a failure to anchor reasoning to a specific agent's knowledge state, rather than from a lack of underlying capability.
We test this hypothesis with a lightweight intervention that asks the model to restate what each character can and cannot observe before answering. The result is a large, consistent improvement that requires no additional parameters.
Given a scenario and a question about agent a, we prompt the model to enumerate the events a witnessed, mark events hidden from a, and only then resolve the query conditioned on a's partial observation. This decomposition mirrors the sub-steps humans perform implicitly.
End-to-end prompting conflates the beliefs of different characters, causing errors that compound as stories grow longer. We argue that an explicit, symbolic representation of "who knows what" is both interpretable and more robust than relying on the model's latent state.
For each character we instantiate a graph whose nodes are propositions and whose edges encode observation events. Queries are answered by reading off the relevant character's sub-graph, sidestepping the interference that plagues monolithic prompting.
LLMs such as GPT-4 seemingly excel at tracking characters’ beliefs in stories, yet struggle to translate this capability into strategic action. The core challenge lies in identifying the implicit inferences about mental states — without being explicitly asked, as in ToMi — that lead to choosing the correct action.
The foresee step prompts the model to anticipate likely future events and the challenges facing each character; the reflect step reasons over candidate actions against those futures before committing. FaR generalizes to diverse out-of-distribution story structures and scenarios.
We use the Sally–Anne unexpected-transfer task and the Smarties unexpected-contents task, each with carefully constructed controls to rule out shallow lexical shortcuts. Items are presented one sentence at a time to probe incremental belief updates.
Emergence of ToM-like behavior may be an unintended by-product of training on human-generated text. We caution against both over-attribution of mental states and premature dismissal, and call for adversarial benchmarks that separate the two.
Every agent edit gets compiled. When the build breaks, the agent opens the logs, reads the error, and fixes its own mistake — then the PDF refreshes, like Overleaf with a mechanic inside.
Large language models increasingly succeed on tasks once thought to require social reasoning, yet their performance on false-belief tasks remains brittle. We hypothesize that much of this brittleness stems from a failure to anchor reasoning to a specific agent's knowledge state, rather than from a lack of underlying capability.
We test this hypothesis with a lightweight intervention that asks the model to restate what each character can and cannot observe before answering. The result is a large, consistent improvement that requires no additional parameters.
Given a scenario and a question about agent a, we prompt the model to enumerate the events a witnessed, mark events hidden from a, and only then resolve the query conditioned on a's partial observation. This decomposition mirrors the sub-steps humans perform implicitly.
End-to-end prompting conflates the beliefs of different characters, causing errors that compound as stories grow longer. We argue that an explicit, symbolic representation of "who knows what" is both interpretable and more robust than relying on the model's latent state.
For each character we instantiate a graph whose nodes are propositions and whose edges encode observation events. Queries are answered by reading off the relevant character's sub-graph, sidestepping the interference that plagues monolithic prompting.
LLMs such as GPT-4 seemingly excel at tracking characters’ beliefs in stories, yet struggle to translate this capability into strategic action. The core challenge lies in identifying the implicit inferences about mental states — without being explicitly asked, as in ToMi — that lead to choosing the correct action.
The foresee step prompts the model to anticipate likely future events and the challenges facing each character; the reflect step reasons over candidate actions against those futures before committing. FaR generalizes to diverse out-of-distribution story structures and scenarios.
We use the Sally–Anne unexpected-transfer task and the Smarties unexpected-contents task, each with carefully constructed controls to rule out shallow lexical shortcuts. Items are presented one sentence at a time to probe incremental belief updates.
Emergence of ToM-like behavior may be an unintended by-product of training on human-generated text. We caution against both over-attribution of mental states and premature dismissal, and call for adversarial benchmarks that separate the two.
Every agent edit gets compiled. When the build breaks, the agent opens the logs, reads the error, and fixes its own mistake — then the PDF refreshes, like Overleaf with a mechanic inside.
Large language models increasingly succeed on tasks once thought to require social reasoning, yet their performance on false-belief tasks remains brittle. We hypothesize that much of this brittleness stems from a failure to anchor reasoning to a specific agent's knowledge state, rather than from a lack of underlying capability.
We test this hypothesis with a lightweight intervention that asks the model to restate what each character can and cannot observe before answering. The result is a large, consistent improvement that requires no additional parameters.
Given a scenario and a question about agent a, we prompt the model to enumerate the events a witnessed, mark events hidden from a, and only then resolve the query conditioned on a's partial observation. This decomposition mirrors the sub-steps humans perform implicitly.
End-to-end prompting conflates the beliefs of different characters, causing errors that compound as stories grow longer. We argue that an explicit, symbolic representation of "who knows what" is both interpretable and more robust than relying on the model's latent state.
For each character we instantiate a graph whose nodes are propositions and whose edges encode observation events. Queries are answered by reading off the relevant character's sub-graph, sidestepping the interference that plagues monolithic prompting.
LLMs such as GPT-4 seemingly excel at tracking characters’ beliefs in stories, yet struggle to translate this capability into strategic action. The core challenge lies in identifying the implicit inferences about mental states — without being explicitly asked, as in ToMi — that lead to choosing the correct action.
The foresee step prompts the model to anticipate likely future events and the challenges facing each character; the reflect step reasons over candidate actions against those futures before committing. FaR generalizes to diverse out-of-distribution story structures and scenarios.
We use the Sally–Anne unexpected-transfer task and the Smarties unexpected-contents task, each with carefully constructed controls to rule out shallow lexical shortcuts. Items are presented one sentence at a time to probe incremental belief updates.
Emergence of ToM-like behavior may be an unintended by-product of training on human-generated text. We caution against both over-attribution of mental states and premature dismissal, and call for adversarial benchmarks that separate the two.
Ask what's missing. The find agent reads your draft, skims your library, and searches with Gemini. You pick the papers — literati.bib updates itself and they're instantly citeable.
Large language models increasingly succeed on tasks once thought to require social reasoning, yet their performance on false-belief tasks remains brittle. We hypothesize that much of this brittleness stems from a failure to anchor reasoning to a specific agent's knowledge state, rather than from a lack of underlying capability.
We test this hypothesis with a lightweight intervention that asks the model to restate what each character can and cannot observe before answering. The result is a large, consistent improvement that requires no additional parameters.
Given a scenario and a question about agent a, we prompt the model to enumerate the events a witnessed, mark events hidden from a, and only then resolve the query conditioned on a's partial observation. This decomposition mirrors the sub-steps humans perform implicitly.
End-to-end prompting conflates the beliefs of different characters, causing errors that compound as stories grow longer. We argue that an explicit, symbolic representation of "who knows what" is both interpretable and more robust than relying on the model's latent state.
For each character we instantiate a graph whose nodes are propositions and whose edges encode observation events. Queries are answered by reading off the relevant character's sub-graph, sidestepping the interference that plagues monolithic prompting.
LLMs such as GPT-4 seemingly excel at tracking characters’ beliefs in stories, yet struggle to translate this capability into strategic action. The core challenge lies in identifying the implicit inferences about mental states — without being explicitly asked, as in ToMi — that lead to choosing the correct action.
The foresee step prompts the model to anticipate likely future events and the challenges facing each character; the reflect step reasons over candidate actions against those futures before committing. FaR generalizes to diverse out-of-distribution story structures and scenarios.
We use the Sally–Anne unexpected-transfer task and the Smarties unexpected-contents task, each with carefully constructed controls to rule out shallow lexical shortcuts. Items are presented one sentence at a time to probe incremental belief updates.
Emergence of ToM-like behavior may be an unintended by-product of training on human-generated text. We caution against both over-attribution of mental states and premature dismissal, and call for adversarial benchmarks that separate the two.
Each item is generated from a causal graph over desires, percepts, and beliefs, letting us hold reasoning structure fixed while varying surface content. This isolates belief tracking from lexical shortcuts and supports large, balanced test sets.
Perspective-taking prompts close much, but not all, of the gap to humans. Robustness varies sharply across scenario types, suggesting current models track beliefs opportunistically rather than systematically.