전체 글 (309) 썸네일형 리스트형 Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning ■ golden-standard Chain-of-Thought (CoT) rationales로 LLM을 fine-tuning하거나, 이를 few-shot prompting에서 correct examples로 사용하는 것이 LLM에 도움이 된다는 것을 보여주는 연구들이 등장했다. ■ 인간은 correct examples을 모방하여 학습할 수도 있지만, 자신의 mistakes로부터 학습하는 것 역시 human cognition의 또 다른 중요한 측면이다. ■ 논문에서는 "LLM은 자신의 mistakes로부터 학습했을 때, reasoning 능력 측면에서 그 mistakes로부터 이점을 얻을 수 있는가?"에 초점을 둔다. 그리고 이 문제를 prompting과 model-tuning이라는 두 관점에서 확인한다... [DeepSeek-V3.2] Pushing the Frontier of Open Large Language Models ■ DeepSeek-V3.2의 key technical breakthroughs는 다음과 같다: (1) efficient attention mechanism인 DeepSeek Sparse Attention (DSA) (2) Scalable Reinforcement Learning Framework (3) Large-Scale Agentic Task Synthesis Pipeline [2512.02556] DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsWe introduce DeepSeek-V3.2, a model tha.. [Self-Error-Instruct] Generalizing from Errors for LLMs Mathematical Reasoning ■ LLM은 mathematical reasoning에서는 여전히 많은 bad cases에서 어려움을 겪는다. 기존의 approaches은 bad case 하나하나에서만 training data를 확장하여 생성하기 때문에, 이러한 bad case들에 내재된 더 넓은 error pattern을 충분히 generalize하지 못한다. ■ 논문에서는 더 generalizable하면서도 targeted된 training data를 합성하기 위한 framework인 Self-Error-Instruct (SEI)를 제안한다. [2505.22591] Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning Self-Error-Ins.. [SDFT] Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning ■ 논문에서는 language model을 downstream task에 fine-tuning하는 과정에서 발생하는 catastrophic forgetting과, fine-tuning 과정에서 발생하는 distribution shift가 general task capability뿐 아니라 model의 safety alignment와 helpfulness까지 저하시킬 수 있음을 보여준다. ■ 저자들은 task dataset과 LLM 사이의 distribution gap이 이러한 문제를 일으키는 주된 근본 원인이라고 가정하며, 이 문제를 해결하기 위한 방법으로 Self-Distillation Fine-Tuning (SDFT)을 제안한다. ■ SDFT는 model 자신이 생성한 distilled datase.. [Self-Instruct] Aligning Language Models with Self-Generated Instructions ■ large "instruction-tuned" language models, 즉 instructions에 응답하도록 fine-tuning된 language models은 새로운 tasks에 zero-shot으로 generalize하는 뛰어난 능력을 보여 왔다. ■ 그럼에도 이러한 model들은 human-written instruction data에 크게 의존하며, 이러한 data는 quantity, diversity, creativity가 제한되어 있는 경우가 있기 때문에 tuning된 model의 generality를 저해한다. ■ 논문에서는 pretrained language model 자신의 generation을 이용해 bootstrapping함으로써 instruction-following c.. Self-Knowledge Distillation in Natural Language Processing ■ deep learning model들의 다양한 NLP tasks에 대한 뛰어난 성능은 deep learning model의 효율적인 knowledge representation 때문인 것으로 볼 수 있다. ■ 더 효율적인 representation을 학습하기 위한 많은 방법이 제안되어 왔지만, pretrained deep network로부터 수행하는 knowledge distillation은 다른 neural network를 학습할 때 soft target probability에 포함된 더 많은 정보를 사용할 수 있음을 보여주었다. ■ 논문에서는 새로운 knowledge distillation method인 "self-knowledge distillation"을 제안한다. ■ 이 방법은 현재 trai.. [CoT-Valve] Length-Compressible Chain-of-Thought Tuning ■ CoT는 모델의 reasoning 능력을 크게 향상시키지만, long chains을 생성하기 때문에 inference 비용도 상당히 증가한다.■ 논문에서 제안하는 CoT-Valve는 모델이 서로 다른 길이의 reasoning chains을 생성할 수 있도록 설계된 새로운 tuning 및 inference strategy이다. ■ 이를 달성하기 위해 parameter space 안에서 특정한 하나의 방향을 찾아내고, 찾아낸 방향을 활용하여 CoT의 길이를 효과적으로 제어한다. ■ 실험 결과, CoT-Valve는 reasoning chain의 길이를 성공적으로 제어하고 압축할 수 있었으며, prompt-based의 길이 제어보다 더 좋은 성능을 보였다. [2502.09601] CoT-Valve: Le.. Chain of Draft: Thinking Faster by Writing Less ■ LLM은 CoT prompting과 같은 step-by-step reasoning 방법을 통해 complex reasoning tasks을 해결하는 데 뛰어난 성능을 보여 왔다. ■ 그러나 인간은 보통 더 효율적인 전략을 사용하는데, 핵심 정보만을 담은 간결한 intermediate thoughts를 draft로 적어 나간다. ■ 논문에서는 인간의 cognitive process에서 영감을 받은 새로운 패러다임으로 Chain of Draft (CoD)를 제안한다. CoD는 LLM이 task를 해결하는 과정에서 최소한이면서도 필요한 정보를 담은 intermediate reasoning outputs을 생성한다. ■ CoD는 verbosity를 줄이고 critical insights에만 집중함으로써, .. 이전 1 2 3 4 ··· 39 다음