Language Models
Reasoning, dialogue, memory, personalization, and robust evaluation.
Associate Professor
Department of Artificial Intelligence
College of Computing
Yonsei University
Office: 신촌캠퍼스 공학원 447
Lab: LangAGI Lab
Email: jinyeo@yonsei.ac.kr
Reasoning, dialogue, memory, personalization, and robust evaluation.
Long-horizon agents, world models, tool use, and interactive learning.
Collaborative decision making, behavior alignment, and clinical agents.
Multimodal clinical reasoning, ophthalmic imaging, and decision support.
COLM 2026
Scalable workflow-grounded environments and reward designs for training procedurally compliant language agents.
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length
ICML 2026 Featured in DAIR.AI (Top AI Paper of the Week)
Studies how task horizon length affects language-model training and performance on long-horizon tasks.
EMBGUARD: Constructing Hazard-Aware Guardrails for Safe Planning in Embodied Agents
ICML 2026
Introduces hazard-aware guardrails for safer planning and action selection in embodied agents.
PAC-BENCH: Evaluating Multi-Agent Collaboration under Privacy Constraints
ACL 2026 Findings
Benchmarks how multiple agents collaborate while respecting privacy constraints.
Why These Documents? Explainable Generative Retrieval with Hierarchical Category Paths
ACL 2026 Findings
Makes generative retrieval more interpretable through hierarchical category-path explanations.
MVIGER: Multi-View Variational Integration of Complementary Knowledge for Generative Recommender
SIGIR 2026
Combines complementary knowledge views through variational integration for generative recommendation.
International Journal of Medical Informatics, 2026 (and Asia Retina Congress 2025) Best Poster Award
Analyzes when human–AI collaboration succeeds or fails on difficult real-world clinical reasoning cases.
ICLR 2026
Examines personalization in embodied agents through the challenges and opportunities of memory utilization.
Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For Perplexity
ICLR 2026
Proposes a fast prior-based method for filtering noisy text as an alternative to perplexity-based selection.
Quantifying Genuine Awareness in Hallucination Prediction: Disentangling Question-Side Shortcuts
Agentic AI in the Wild Workshop, ICLR 2026
Separates genuine hallucination awareness from shortcuts that rely only on properties of the question.
AgenticShop: Benchmarking Agentic Product Curation for Personalized Web Shopping
WWW 2026
Introduces a benchmark for personalized product curation by web-shopping agents.
Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
NeurIPS 2025 Spotlight
Develops process reward models that better reinforce web agents during multistep interaction.
Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuning
NeurIPS 2025 Spotlight
Improves the speed and fluency of diffusion language models with convolutional decoding and rejective fine-tuning.
ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions
EMNLP 2025 Findings
Stress-tests tool-augmented language models in realistic, long-term tool-use interactions.
PRINCIPLES: Synthetic Strategy Memory for Proactive Dialogue Agents
EMNLP 2025 Findings
Uses synthetic strategy memories to make dialogue agents more proactive and consistent.
EMNLP 2025 Findings
Evaluates conversational recommenders with target-free user simulation rather than hidden target guessing.
EMNLP 2025 Findings
Investigates whether English–Korean code-switching activates distinct knowledge behavior in language models.
Designing Memory-Augmented AR Agents for Spatiotemporal Reasoning in Personalized Task Assistance
XRMemory Workshop, IEEE ISMAR 2025
Designs memory-augmented augmented-reality agents for personalized spatiotemporal task assistance.
Rethinking Reward Model Evaluation Through the Lens of Reward Overoptimization
ACL 2025
Reframes reward-model evaluation around robustness to reward overoptimization.
ACL 2025
Benchmarks whether large language models can understand and generate structured scene graphs.
Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
ACL 2025 Findings
Models client metacognition and openness to improve evaluation of LLM-based therapeutic conversations.
ACL 2025 Industry
Provides data for mitigating the cold-start problem when training short chain-of-thought reasoning models with reinforcement learning.
Ophthalmology Science, 2025
Studies age-related hypofluorescent spots as a prognostic factor in polypoidal choroidal vasculopathy.
Review-driven Personalized Preference Reasoning with Large Language Models for Recommendation
SIGIR 2025
Uses review evidence to improve personalized preference reasoning for recommendation.
Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web Navigation
ICLR 2025
Learns and exploits environment dynamics to improve long-horizon web navigation.
Towards Lifelong Dialogue Agents via Timeline-based Memory Management
NAACL 2025
Manages conversational memories along a timeline to support lifelong dialogue agents.
NAACL 2025 Findings
Introduces a psychometric test set for measuring personality consistency in large language models.
Nature Scientific Reports, 2025
Predicts future branch retinal vein occlusion from multimodal pre-onset fundus images.
WSDM 2025
Aligns cross-lingual entities without supervision by matching neighboring triples and textual information.
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
NeurIPS 2024
Meta-learns where a model should focus when continually acquiring new knowledge.
EMNLP 2024 Featured in Hugging Face (Daily Papers)
Uses pseudocode execution as an intermediate reasoning process for algorithmic problem solving.
COFEE-GYM: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous Code
EMNLP 2024
Provides an environment for evaluating and improving natural-language feedback on erroneous code.
Evidence-Focused Fact Summarization for Knowledge-Augmented Zero-shot Question Answering
EMNLP 2024
Summarizes evidence relevant to knowledge-augmented zero-shot question answering.
CACTUS: Towards Psychological Counseling Conversations using Cognitive Behavioral Theory
EMNLP 2024 Findings
Builds psychologically grounded counseling conversations using principles from cognitive behavioral theory.
EMNLP 2024 Findings
Uses question-driven reasoning to identify implicit knowledge for more insightful table summaries.
ACL 2024 Outstanding Paper Award
Reduces preference bias to improve emotional-support conversations with large language models.
VERIFINER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language Models
ACL 2024
Augments named-entity recognition with knowledge-grounded reasoning and verification.
PEARL: A Review-driven Persona-Knowledge grounded Conversational Recommendation Dataset
ACL 2024 Findings
Introduces a review-driven persona- and knowledge-grounded dataset for conversational recommendation.
Self-Consistent Reasoning-based Aspect-Sentiment Quad Prediction with Extract-then-Assign Strategy
ACL 2024 Findings
Applies self-consistent extract-then-assign reasoning to aspect–sentiment quad prediction.
RTSUM: Relation Triple-based Interpretable Summarization with Multi-level Salience Visualization
NAACL 2024 Demonstration
Uses question-driven reasoning to identify implicit knowledge for more insightful table summaries.
EACL 2024 Best Paper Award in KCC 2023 (Preliminary study)
Refines persona information with commonsense to improve long-term conversational memory.
Evidentiality-Aware Retrieval for Overcoming Abstractiveness in Open-Domain Question Answering
EACL 2024 Findings
Uses evidentiality-aware retrieval to address abstractiveness in open-domain question answering.
Investigative Ophthalmology & Visual Science (IOVS), 2024
Uses multitask learning to distinguish necrotizing viral from non-infectious retinitis using blood and serology data.
AAAI 2024
Introduces reasoning-aware diagnostic prompting with model-generated clinical rationales.
Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational Agents
EMNLP 2023
Distills dialogue chain-of-thought reasoning into commonsense-aware conversational agents.
European Radiology, 2023
Predicts glioma IDH genotype from free-text MR radiology reports using natural language processing.
CoTEVer: Chain of Thought Prompting Annotation Toolkit for Explanation Verification
EACL 2023 Demonstration
Provides an annotation toolkit for verifying chain-of-thought explanations.
Evidence-Empowered Transfer Learning for Alzheimer's Disease
IEEE International Symposium on Biomedical Imaging (ISBI), 2023
Transfers evidence across tasks to improve learning for Alzheimer’s disease prediction.
TUTORING: Instruction-grounded Conversational Agent for Language Learners
AAAI 2023 Demonstration
Builds an instruction-grounded conversational tutor for language learners.
EMNLP 2022
Automatically curates large-scale multi-skill dialogue datasets from machine-generated conversations.
Mind the Gap! Injecting Commonsense Knowledge for Abstractive Dialogue Summarization
COLING 2022
Injects commonsense knowledge to improve abstractive dialogue summarization.
Modularized Transfer Learning with Multiple Knowledge Graphs for Zero-shot Commonsense Reasoning
NAACL 2022
Transfers knowledge from multiple graphs for zero-shot commonsense reasoning.
Dual Task Framework for Improving Persona-grounded Dialogue Dataset
AAAI 2022
Improves persona-grounded dialogue data with a dual-task learning framework.
TrustAL: Trustworthy Active Learning using Knowledge Distillation
AAAI 2022
Combines active learning and knowledge distillation for more trustworthy sample selection.
Nature Scientific Reports, 2022
Predicts photodynamic-therapy outcomes for chronic central serous chorioretinopathy using multimodal transfer learning.
Label and Context Augmentation for Response Selection at DSTC8
IEEE/ACM TASLP, 2021
Augments labels and conversational context to improve response selection.
Less is More: Attention Supervision with Counterfactuals for Text Classification
EMNLP 2020
Uses counterfactual attention supervision to improve text classification with limited annotation.
Conversion Prediction from Clickstream: Modeling Market Prediction and Customer Predictability
IEEE TKDE, 2020
Models market-level and customer-level predictability for online purchase conversion.
XINA: Explainable Instance Alignment using Dominance Relationship
IEEE TKDE, 2020
Explains instance alignment through dominance relationships between candidate matches.
Learning with Limited Data for Multilingual Reading Comprehension
EMNLP 2019
Transfers limited supervision across languages for multilingual reading comprehension.
Soft Representation Learning for Sparse Transfer
ACL 2019
Learns soft representations that improve transfer under sparse supervision.
Translations as Additional Contexts for Sentence Classification
IJCAI 2018
Uses machine translations as auxiliary contexts for sentence classification.
Visual Choice of Plausible Alternatives: An Evaluation of Image-based Commonsense Causal Reasoning
LREC 2018
Evaluates image-based commonsense causal reasoning through plausible alternative selection.
Machine-translated Knowledge Transfer for Commonsense Causal Reasoning
AAAI 2018
Transfers machine-translated knowledge for commonsense causal reasoning.
Efficient Keyword-aware Representative Travel Route Recommendation
IEEE TKDE, 2017
Recommends representative travel routes while accounting for user-specified keywords.
Multimodal KB Harvesting for Emerging Spatial Entities
IEEE TKDE, 2017
Harvests knowledge about emerging spatial entities from multimodal sources.
Predicting Online Purchase Conversion for Retargeting
WSDM 2017
Predicts purchase conversion to support personalized retargeting.
Event Grounding from Multimodal Social Network Fusion
ICDM 2016
Grounds events by fusing multimodal evidence from social networks.
Browsing2purchase: Online Customer Model for Sales Forecasting in an E-Commerce Site
WWW 2016
Models browsing behavior to forecast purchases in e-commerce.
Purchase Influence Mining: Identifying Top-k Items Attracting Purchase of Target Item
WWW 2016
Identifies products that most strongly influence purchases of a target item.
Understanding Emerging Spatial Entities
AAAI 2016
Models and understands newly emerging spatial entities from web data.
KSTR: Keyword-aware Skyline Travel Route Recommendation
ICDM 2015
Recommends keyword-aware skyline travel routes under multiple user preferences.
Finding Influential Products on Social Domination Game
CIKM 2012
Identifies influential products using a social-domination formulation.
I lead the LangAGI Lab at Yonsei University. We study language models, agents, interaction, and medical AI. Prospective students may contact me with a CV and a brief description of their research interests.
I am an Associate Professor in the Department of Artificial Intelligence at Yonsei University. Before joining Yonsei University, I worked at SKT Brain. I received my Ph.D. from POSTECH and completed research internships at Adobe Research, San Jose.