{"ok":true,"capturedAt":"2026-09-25T04:01:08.053Z","venues":["ACL 2025","EMNLP 2025","NAACL 2025"],"paper_count":90,"papers":[{"title":"CLEME2.0: Towards Interpretable Evaluation by Disentangling Edits for Grammatical Error Correction","authors":[],"venue_group":"ACL 2025","abstract_snippet":"The paper focuses on the interpretability of Grammatical Error Correction (GEC) evaluation metrics, which received little attention in previous studies. To bridge the gap, we introduce CLEME2.0, a reference-based metric describing four f...","url":"https://github.com/THUKElab/CLEME","doi":"10.18653/v1/2025.acl-long.10"},{"title":"LLM-Powered Test Case Generation for Detecting Bugs in Plausible Programs","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Detecting tricky bugs in plausible programs, those that pass existing test suites yet still contain bugs, remains a significant challenge in software testing. To address this problem, we propose TrickCatcher, an LLM-powered approach to g...","url":"https://github.com/RinCloud/TrickCatcher/","doi":"10.18653/v1/2025.acl-long.20"},{"title":"BelarusianGLUE: Towards a Natural Language Understanding Benchmark for Belarusian","authors":[],"venue_group":"ACL 2025","abstract_snippet":"In the epoch of multilingual large language models (LLMs), it is still challenging to evaluate the models’ understanding of lower-resourced languages, which motivates further development of expert-crafted natural language understanding b...","url":"https://hf.co/datasets/maaxap/BelarusianGLUE","doi":"10.18653/v1/2025.acl-long.25"},{"title":"RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios","authors":[],"venue_group":"ACL 2025","abstract_snippet":"This paper introduces RuleArena, a novel and challenging benchmark designed to evaluate the ability of large language models (LLMs) to follow complex, real-world rules in reasoning. Covering three practical domains – airline baggage fees...","url":"https://github.com/skyriver-2000/rulearena","doi":"10.18653/v1/2025.acl-long.27"},{"title":"Semantic Exploration with Adaptive Gating for Efficient Problem Solving with Language Models","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Recent advancements in large language models (LLMs) have shown remarkable potential in various complex tasks requiring multi-step reasoning methods like tree search to explore diverse reasoning paths. However, existing methods often suff...","url":"https://github.com/ml-postech/SEAG-semantic-exploration-with-adaptive-gating","doi":"10.18653/v1/2025.acl-long.29"},{"title":"Can Multimodal Large Language Models Understand Spatial Relations?","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Spatial relation reasoning is a crucial task for multimodal large language models (MLLMs) to understand the objective world. However, current benchmarks have issues like relying on bounding boxes, ignoring perspective substitutions, or a...","url":"https://huggingface.co/datasets/liuziyan/SpatialMQA","doi":"10.18653/v1/2025.acl-long.31"},{"title":"BPP-Search: Enhancing Tree of Thought Reasoning for Mathematical Modeling Problem Solving","authors":[],"venue_group":"ACL 2025","abstract_snippet":"LLMs exhibit advanced reasoning capabilities, offering the potential to transform natural language questions into mathematical models. However, existing open-source datasets in operations research domain lack detailed annotations of the...","url":"https://huggingface.co/datasets/LLM4OR/StructuredOR","doi":"10.18653/v1/2025.acl-long.40"},{"title":"LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Large vision language models (LVLMs) have improved the document understanding capabilities remarkably, enabling the handling of complex document elements, longer contexts, and a wider range of tasks. However, existing document understand...","url":"https://github.com/dengc2023/LongDocURL","doi":"10.18653/v1/2025.acl-long.57"},{"title":"Modeling Uncertainty in Composed Image Retrieval via Probabilistic Embeddings","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Composed Image Retrieval (CIR) enables users to search for images using multimodal queries that combine text and reference images. While metric learning methods have shown promise, they rely on deterministic point embeddings that fail to...","url":"https://github.com/tanghme0w/ACL25-CoPE","doi":"10.18653/v1/2025.acl-long.61"},{"title":"APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model Prompts","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Large Language Models (LLMs) have become increasingly capable of handling diverse tasks with the aid of well-crafted prompts and integration of external tools, but as task complexity rises, the workflow involving LLMs can be complicated...","url":"https://github.com/appl-team/appl","doi":"10.18653/v1/2025.acl-long.63"},{"title":"Autoregressive Speech Synthesis without Vector Quantization","authors":[],"venue_group":"ACL 2025","abstract_snippet":"We present MELLE, a novel continuous-valued token based language modeling approach for text-to-speech synthesis (TTS). MELLE autoregressively generates continuous mel-spectrogram frames directly from text condition, bypassing the need fo...","url":"https://aka.ms/melle","doi":"10.18653/v1/2025.acl-long.65"},{"title":"“Yes, My LoRD.” Guiding Language Model Extraction with Locality Reinforced Distillation","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Model extraction attacks (MEAs) on large language models (LLMs) have received increasing attention in recent research. However, existing attack methods typically adapt the extraction strategies originally developed for deep neural networ...","url":"https://github.com/liangzid/LoRD-MEA","doi":"10.18653/v1/2025.acl-long.73"},{"title":"Jailbreak Large Vision-Language Models Through Multi-Modal Linkage","authors":[],"venue_group":"ACL 2025","abstract_snippet":"With the rapid advancement of Large Vision-Language Models (VLMs), concerns about their ‌potential misuse and abuse have grown rapidly. Prior research has exposed VLMs’ vulnerability to jailbreak attacks, where carefully crafted inputs c...","url":"https://github.com/wangyu-ovo/MML","doi":"10.18653/v1/2025.acl-long.74"},{"title":"MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset","authors":[],"venue_group":"ACL 2025","abstract_snippet":"To enable Large Language Models (LLMs) to function as conscious agents with generalizable reasoning capabilities, it is crucial that they possess the ability to comprehend situational changes (transitions) in distribution triggered by en...","url":"https://github.com/HKUST-KnowComp/MARS","doi":"10.18653/v1/2025.acl-long.79"},{"title":"Disentangling Memory and Reasoning Ability in Large Language Models","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Large Language Models (LLMs) have demonstrated strong performance in handling complex tasks that require both extensive knowledge and reasoning abilities. However, the existing LLM inference pipeline operates as an opaque process without...","url":"https://github.com/MingyuJ666/Disentangling-Memory-and-Reasoning","doi":"10.18653/v1/2025.acl-long.84"},{"title":"Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attributions Explainability","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Deep neural network predictions are notoriously difficult to interpret. Feature attribution methods aim to explain these predictions by identifying the contribution of each input feature. Faithfulness, often evaluated using the area over...","url":"https://github.com/JoakimEdin/naopc","doi":"10.18653/v1/2025.acl-long.86"},{"title":"LangSAMP: Language-Script Aware Multilingual Pretraining","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Recent multilingual pretrained language models (mPLMs) often avoid using language embeddings – learnable vectors assigned to individual languages. However, this places a significant burden on token representations to encode all language-...","url":"https://github.com/cisnlp/LangSAMP","doi":"10.18653/v1/2025.acl-long.88"},{"title":"RelationalCoder: Rethinking Complex Tables via Programmatic Relational Transformation","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Semi-structured tables, with their varied layouts and formatting artifacts, remain a major obstacle for automated data processing and analytics. To address these challenges, we propose RelationalCoder, which uniformly converts semi-struc...","url":"https://github.com/haoyudong/RelationalCoder","doi":"10.18653/v1/2025.acl-long.89"},{"title":"From Information to Insight: Leveraging LLMs for Open Aspect-Based Educational Summarization","authors":[],"venue_group":"ACL 2025","abstract_snippet":"This paper addresses the challenge of aspect-based summarization in education by introducing Reflective ASPect-based summarization (ReflectASP), a novel dataset that summarizes student reflections on STEM lectures. Despite the promising...","url":"https://github.com/cs329yangzhong/ReflectASP","doi":"10.18653/v1/2025.acl-long.95"},{"title":"Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion","authors":[],"venue_group":"ACL 2025","abstract_snippet":"This paper addresses the critical need for democratizing large language models (LLM) in the Arab world, a region that has seen slower progress in developing models comparable to state-of-the-art offerings like GPT-4 or GPT-3.5, due to a...","url":"https://github.com/FreedomIntelligence/AraLLaMa","doi":"10.18653/v1/2025.acl-long.100"},{"title":"CompileAgent: Automated Real-World Repo-Level Compilation with Tool-Integrated LLM-based Agent System","authors":[],"venue_group":"ACL 2025","abstract_snippet":"With open-source projects growing in size and complexity, manual compilation becomes tedious and error-prone, highlighting the need for automation to improve efficiency and accuracy. However, the complexity of compilation instruction sea...","url":"https://github.com/Ch3nYe/AutoCompiler","doi":"10.18653/v1/2025.acl-long.103"},{"title":"AndroidLab: Training and Systematic Benchmarking of Android Autonomous Agents","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Autonomous agents have become increasingly important for interacting with the real world. Android agents, in particular, have been a frequently-mentioned interaction method. However, existing studies for training and evaluating Android a...","url":"https://github.com/THUDM/Android-Lab","doi":"10.18653/v1/2025.acl-long.107"},{"title":"Multimodal Transformers are Hierarchical Modal-wise Heterogeneous Graphs","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Multimodal Sentiment Analysis (MSA) is a rapidly developing field that integrates multimodal information to recognize sentiments, and existing models have made significant progress in this area. The central challenge in MSA is multimodal...","url":"https://github.com/drewjin/GsiT.git","doi":"10.18653/v1/2025.acl-long.109"},{"title":"Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Nonverbal communication is integral to human interaction, with gestures, facial expressions, and body language conveying critical aspects of intent and emotion. However, existing large language models (LLMs) fail to effectively incorpora...","url":"https://github.com/winston1214/nonverbal-conversation","doi":"10.18653/v1/2025.acl-long.112"},{"title":"LegalAgentBench: Evaluating LLM Agents in Legal Domain","authors":[],"venue_group":"ACL 2025","abstract_snippet":"With the increasing intelligence and autonomy of LLM Agents, their potential applications in the legal domain are becoming increasingly apparent. However, existing general-domain benchmarks are unable to fully capture the complexity and...","url":"https://github.com/CSHaitao/LegalAgentBench","doi":"10.18653/v1/2025.acl-long.116"},{"title":"Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Recent English Common Crawl datasets like FineWeb-Edu and DCLM achieved significant benchmark gains via aggressive model-based filtering, but at the cost of removing 90% of data. This limits their suitability for long token horizon train...","url":"https://data.commoncrawl.org/contrib/Nemotron/Nemotron-CC/index.html","doi":"10.18653/v1/2025.acl-long.123"},{"title":"Are Any-to-Any Models More Consistent Across Modality Transfers Than Specialists?","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Any-to-any generative models aim to enable seamless interpretation and generation across multiple modalities within a unified framework, yet their ability to preserve relationships across modalities remains uncertain. Do unified models t...","url":"https://github.com/JiwanChung/ACON","doi":"10.18653/v1/2025.acl-long.130"},{"title":"AndroidGen: Building an Android Language Agent under Data Scarcity","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Large language models have opened up a world of possibilities for various NLP tasks, sparking optimism for the future. Despite their potential, LLMs have yet to be widely used as agents on real mobile devices. The main challenge is the n...","url":"https://github.com/THUDM/AndroidGen","doi":"10.18653/v1/2025.acl-long.138"},{"title":"Prompt Candidates, then Distill: A Teacher-Student Framework for LLM-driven Data Annotation","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Recently, Large Language Models (LLMs) have demonstrated significant potential for data annotation, markedly reducing the labor costs associated with downstream applications. However, existing methods mostly adopt an aggressive strategy...","url":"https://github.com/MingxuanXia/CanDist","doi":"10.18653/v1/2025.acl-long.139"},{"title":"ToolHop: A Query-Driven Benchmark for Evaluating Large Language Models in Multi-Hop Tool Use","authors":[],"venue_group":"ACL 2025","abstract_snippet":"Effective evaluation of multi-hop tool use is critical for analyzing the understanding, reasoning, and function-calling capabilities of large language models (LLMs). However, progress has been hindered by a lack of reliable evaluation da...","url":"https://huggingface.co/datasets/bytedance-research/ToolHop","doi":"10.18653/v1/2025.acl-long.150"},{"title":"Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Although Contrastive Language-Image Pre-training (CLIP) exhibits strong performance across diverse vision tasks, its application to person representation learning faces two critical challenges: (i) the scarcity of large-scale annotated v...","url":"https://github.com/Multimodal-Representation-Learning-MRL/GA-DMS","doi":"10.18653/v1/2025.emnlp-main.14"},{"title":"ToneCraft: Cantonese Lyrics Generation with Harmony of Tones and Pitches","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Lyrics generation has garnered increasing attention within the artificial intelligence community. Our task focuses on generating harmonious Cantonese lyrics. Unlike other languages, Cantonese has a unique system of nine contours and six...","url":"https://github.com/purepasser-by/ToneCraft","doi":"10.18653/v1/2025.emnlp-main.18"},{"title":"SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"We introduce SensorLLM, a two-stage framework that enables Large Language Models (LLMs) to perform human activity recognition (HAR) from sensor time-series data. Despite their strong reasoning and generalization capabilities, LLMs remain...","url":"https://github.com/zechenli03/SensorLLM","doi":"10.18653/v1/2025.emnlp-main.19"},{"title":"DeepResearcher: Scaling Deep Research via Reinforcement Learning in Real-world Environments","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Large Language Models (LLMs) with web search capabilities show significant potential for deep research, yet current methods—brittle prompt engineering or RAG-based reinforcement learning in controlled environments—fail to capture real-wo...","url":"https://github.com/GAIR-NLP/DeepResearcher","doi":"10.18653/v1/2025.emnlp-main.22"},{"title":"SoundMind: RL-Incentivized Logic Reasoning for Audio-Language Models","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"While large language models have demonstrated impressive reasoning abilities, their extension to the audio modality, particularly within large audio-language models (LALMs), remains underexplored. Addressing this gap requires a systemati...","url":"https://github.com/xid32/SoundMind","doi":"10.18653/v1/2025.emnlp-main.27"},{"title":"CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Chain-of-Thought (CoT) reasoning enhances Large Language Models (LLMs) by encouraging step-by-step reasoning in natural language. However, leveraging a latent continuous space for reasoning may offer benefits in terms of both efficiency...","url":"https://github.com/zhenyi4/codi","doi":"10.18653/v1/2025.emnlp-main.36"},{"title":"Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style. Driven by rising industrial demand and breakthroughs in deep learning, e.g., diff...","url":"https://github.com/imxtx/awesome-controllabe-speech-synthesis","doi":"10.18653/v1/2025.emnlp-main.40"},{"title":"EMNLP: Educator-role Moral and Normative Large Language Models Profiling","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Simulating Professions (SP) enables Large Language Models (LLMs) to emulate professional roles. However, comprehensive psychological and ethical evaluation in these contexts remains lacking. This paper introduces EMNLP, an Educator-role...","url":"https://e-m-n-l-p.github.io/","doi":"10.18653/v1/2025.emnlp-main.42"},{"title":"TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical Domain","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"While document summarization with LLMs has enhanced access to textual information, concerns about the factual accuracy of these summaries persist (e.g., hallucination), especially in the medical domain. Tracing source evidence from which...","url":"https://github.com/chubohao/TracSum","doi":"10.18653/v1/2025.emnlp-main.43"},{"title":"Parallel Continuous Chain-of-Thought with Jacobi Iteration","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Continuous chain-of-thought has been shown to be effective in saving reasoning tokens for large language models. By reasoning with continuous latent thought tokens, continuous CoT is able to perform implicit reasoning in a compact manner...","url":"https://github.com/whyNLP/PCCoT","doi":"10.18653/v1/2025.emnlp-main.47"},{"title":"LinkAlign: Scalable Schema Linking for Real-World Large-Scale Multi-Database Text-to-SQL","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Schema linking is a critical bottleneck in applying existing Text-to-SQL models to real-world, large-scale, multi-database environments. Through error analysis, we identify two major challenges in schema linking: (1) Database Retrieval:...","url":"https://github.com/Satissss/LinkAlign","doi":"10.18653/v1/2025.emnlp-main.51"},{"title":"On Relation-Specific Neurons in Large Language Models","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"In large language models (LLMs), certain neurons can store distinct pieces of knowledge learned during pretraining. While factual knowledge typically appears as a combination of relations and entities, it remains unclear whether some neu...","url":"https://github.com/cisnlp/relation-specific-neurons","doi":"10.18653/v1/2025.emnlp-main.52"},{"title":"Weight-Aware Activation Sparsity with Constrained Bayesian Optimization Scheduling for Large Language Models","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Activation sparsity provides a dynamic, input-dependent alternative to weight pruning for accelerating inference in large language models (LLMs), effectively reducing unnecessary computations and memory accesses during the forward pass....","url":"https://github.com/HITSZ-Miao-Group/WAS","doi":"10.18653/v1/2025.emnlp-main.57"},{"title":"VC4VG: Optimizing Video Captions for Text-to-Video Generation","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Recent advances in text-to-video (T2V) generation highlight the critical role of high-quality video-text pairs in training models capable of producing coherent and instruction-aligned videos. However, strategies for optimizing video capt...","url":"https://github.com/qyr0403/VC4VG","doi":"10.18653/v1/2025.emnlp-main.59"},{"title":"Skeletons Matter: Dynamic Data Augmentation for Text-to-Query","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"The task of translating natural language questions into query languages has long been a central focus in semantic parsing. Recent advancements in Large Language Models (LLMs) have significantly accelerated progress in this field. However...","url":"https://github.com/jjjycaptain/Skeletron","doi":"10.18653/v1/2025.emnlp-main.64"},{"title":"MovieCORE: COgnitive REasoning in Movies","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"This paper introduces MovieCORE, a novel video question answering (VQA) dataset designed to probe deeper cognitive understanding of movie content. Unlike existing datasets that focus on surface-level comprehension, MovieCORE emphasizes q...","url":"https://joslefaure.github.io/assets/html/moviecore.html","doi":"10.18653/v1/2025.emnlp-main.66"},{"title":"Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Reasoning large language models (LLMs) excel in complex tasks, which has drawn significant attention to reinforcement learning (RL) for LLMs. However, existing approaches allocate an equal number of rollouts to all questions during the R...","url":"https://anonymous.4open.science/r/E3-RL4LLMs-DB28","doi":"10.18653/v1/2025.emnlp-main.75"},{"title":"Persuasion Dynamics in LLMs: Investigating Robustness and Adaptability in Knowledge and Safety with DuET-PD","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Large Language Models (LLMs) can struggle to balance gullibility to misinformation and resistance to valid corrections in persuasive dialogues, a critical challenge for reliable deployment. We introduce DuET-PD (Dual Evaluation for Trust...","url":"https://github.com/Social-AI-Studio/DuET-PD","doi":"10.18653/v1/2025.emnlp-main.81"},{"title":"CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMs","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Improvements in model construction, including fortified safety guardrails, allow Large language models (LLMs) to increasingly pass standard safety checks. However, LLMs sometimes slip into revealing harmful behavior, such as expressing r...","url":"https://github.com/nafisenik/CoBia","doi":"10.18653/v1/2025.emnlp-main.84"},{"title":"Beyond the Surface: Measuring Self-Preference in LLM Judgments","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Recent studies show that large language models (LLMs) exhibit self-preference bias when serving as judges, meaning they tend to favor their own responses over those generated by other models. Existing methods typically measure this bias...","url":"https://github.com/zhiyuanc2001/self-preference","doi":"10.18653/v1/2025.emnlp-main.86"},{"title":"Unstructured Evidence Attribution for Long Context Query Focused Summarization","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Large language models (LLMs) are capable of generating coherent summaries from very long contexts given a user query, and extracting and citing evidence spans helps improve the trustworthiness of these summaries. Whereas previous work ha...","url":"https://github.com/dwright37/unstructured-evidence-sunset","doi":"10.18653/v1/2025.emnlp-main.95"},{"title":"RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Multimodal question answering (QA) often requires identifying which video, audio, or sensor tokens are relevant to the question. Yet modality disagreements are common: off-camera speech, background noise, or motion outside the field of v...","url":"https://github.com/BASHLab/RAVEN","doi":"10.18653/v1/2025.emnlp-main.96"},{"title":"Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Vision Language Models (VLMs) have achieved remarkable success in a wide range of vision applications of increasing complexity and scales, yet choosing the right VLM model size involves a trade-off between response quality and cost. Whil...","url":"https://github.com/UIUC-MONET/Cache-of-Thoughts","doi":"10.18653/v1/2025.emnlp-main.97"},{"title":"Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Video large language models (VideoLLM) excel at video understanding, but face efficiency challenges due to the quadratic complexity of abundant visual tokens. Our systematic analysis of token compression methods for VideoLLMs reveals two...","url":"https://github.com/xuyang-liu16/VidCom2","doi":"10.18653/v1/2025.emnlp-main.98"},{"title":"Foot-In-The-Door: A Multi-turn Jailbreak for LLMs","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Ensuring AI safety is crucial as large language models become increasingly integrated into real-world applications. A key challenge is jailbreak, where adversarial prompts bypass built-in safeguards to elicit harmful disallowed outputs....","url":"https://github.com/Jinxiaolong1129/Foot-in-the-door-Jailbreak","doi":"10.18653/v1/2025.emnlp-main.100"},{"title":"F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"With the growing adoption of large language models (LLMs) in NLP tasks, concerns about their fairness have intensified. Yet, most existing fairness benchmarks rely on closed-ended evaluation formats, which diverge from real-world open-en...","url":"https://github.com/VelikayaScarlet/F2Bench","doi":"10.18653/v1/2025.emnlp-main.105"},{"title":"CodeMixBench: Evaluating Code-Mixing Capabilities of LLMs Across 18 Languages","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Code-mixing, the practice of switching between languages within a conversation, poses unique challenges for traditional NLP. Existing benchmarks like LinCE and GLUECoS are limited by their narrow language pairs and tasks, failing to adeq...","url":"https://github.com/Jeromeyluck/CodeMixBench","doi":"10.18653/v1/2025.emnlp-main.109"},{"title":"SafeScientist: Enhancing AI Scientist Safety for Risk-Aware Scientific Discovery","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"Recent advancements in large language model (LLM) agents have significantly accelerated scientific discovery automation, yet concurrently raised critical ethical and safety concerns. To systematically address these challenges, we introdu...","url":"https://github.com/ulab-uiuc/SafeScientist","doi":"10.18653/v1/2025.emnlp-main.116"},{"title":"RuCCoD: Towards Automated ICD Coding in Russian","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"This study investigates the feasibility of automating clinical coding in Russian, a language with limited biomedical resources. We present a new dataset for ICD coding, which includes diagnosis fields from electronic health records (EHRs...","url":"https://github.com/auto-icd-coding/ruccod","doi":"10.18653/v1/2025.emnlp-main.129"},{"title":"VoiceCraft-X: Unifying Multilingual, Voice-Cloning Speech Synthesis and Speech Editing","authors":[],"venue_group":"EMNLP 2025","abstract_snippet":"We introduce VoiceCraft-X, an autoregressive neural codec language model which unifies multilingual speech editing and zero-shot text-to-speech (TTS) synthesis across 11 languages: English, Mandarin, Korean, Japanese, Spanish, French, Ge...","url":"https://zhishengzheng.com/voicecraft-x/","doi":"10.18653/v1/2025.emnlp-main.137"}],"notes":["Recent papers from the current ACL, EMNLP, and NAACL proceedings via the ACL Anthology. A bounded per-venue sample (the first papers in program order); follow url for the full paper. Abstracts (CC-BY) are clipped. The venue list is curated and refreshed each cycle."],"source":{"name":"ACL Anthology","url":"https://aclanthology.org","license":"ACL Anthology metadata; abstracts CC-BY. TensorFeed links and summarizes with a clipped abstract; full papers are linked, not republished."}}