Woojung Song

Integrated M.S.-Ph.D. in Data Science, Seoul National University
Advised by Yohan Jo · HOLI Lab (Human-Oriented Language Intelligence)

Research Interests

I study how language agents behave in the world and how they can model human behavior. I believe that building better agents requires understanding the decisions they make, the environments they operate in, and the people they interact with. My research spans two connected directions: Agents and Social Simulation.

  1. Agents. I'm interested in how agents get things done in the real world. I study how they use tools, work with users, and respond when things don't go as expected. I look at the decisions they make along the way to understand where they struggle and how to make them more reliable.
  2. Social Simulation. I'm also interested in building agents that can take on different roles, from everyday users to characters in a novel. For role-playing language agents, this means staying faithful to characters as their stories unfold and their perspectives evolve. Building on my work on value measurement and behavioral evaluation, I aim to develop social agents that reflect the diversity of human behavior and how it changes across contexts and over time.

I'm currently seeking a summer research internship in agents or social simulation, and I welcome collaborations in these areas. Feel free to reach out at opusdeisong@snu.ac.kr.

News

Recent Publications

See all on the Publications page.

Under Review

AgentHabit: Characterizing Distinct Behaviors of Agents on Everyday Tasks

Woojung Song*, Hoyeol Yang*, Jeonghoon Shim, Sungjib Lim, Jonggeun Lee, Yunho Choi, Yohan Jo  First

Agents can complete the same task while behaving very differently. AgentHabit profiles these differences along 23 behavioral axes across 86 everyday tasks, including when agents ask questions, use tools, and explain their decisions. Across 18 models, these profiles remain recognizable on different task sets, while prompting and fine-tuning change some tendencies more easily than others.

AgentHabit compares two agents' interactions and behavioral profiles on the same everyday task

Under Review

Agents' Overreliance on Unreliable Tools

Hoyeol Yang*, Woojung Song*, Taewon Kim, Jonghyun Song, Seoyeon Park, Yohan Jo  First

Across 14 models, agents frequently adopt incorrect outputs from web search, an LLM sub-agent, and a code executor, sometimes overriding answers they already know. Even when their reasoning notices a conflict, they often pass the incorrect content to users without warning. Prompts that ask agents to verify tool outputs reduce this reliance across all three tools while largely preserving accuracy when the outputs are correct.

Web-search evaluation design

EMNLP 2026Main

ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?

Woojung Song*, Nalim Kim*, Sangjun Song, Chaewon Heo, Jongwon Lim, Yohan Jo  First

Role-playing agents should evolve with their character, not hold a fixed persona. ArcANE, an automatically built benchmark of 17 novels and 80 characters, shows that conditioning on a Character Arc, the narrative segmented into psychological phases, outperforms every other context strategy, most of all on scenarios the source text never explores.

ArcANE construction pipeline

EMNLP 2026Main

Human Psychometric Questionnaires Mischaracterize LLM Behavior

Woojung Song*, Dongmin Choi*, Yoonah Park, Jongwook Han, Eun-Ju Lee, Yohan Jo  First

The psychological profile an LLM reports on a questionnaire is not the one it shows when actually generating text. Familiar questionnaire items cue socially desirable answers, and persona effects that appear on questionnaires vanish on realistic user queries. We argue models should be profiled by what they generate, not what they self-report.

Questionnaire vs generation behavior