본문 바로가기

전체 글

(312)
LoRA Without Regret LoRA Without Regret - Thinking Machines Lab LoRA Without RegretHow LoRA matches full training performance more broadly than expected.thinkingmachines.ai ■ base model의 성능은 model parameter나 data scale이 커질수록 계속해서 향상되고 있다. pre-training stage에서 written-down human knowledge에 존재하는 모든 patterns을 학습하고 표현하기 위해서는 trillions 규모의 파라미터와 토큰이 필요하기 때문이다. ■ post-training은 이와 대조적으로 훨씬 작은 데이터셋을 사용하며, 일반적으로 더 좁은 범위의 ..
LoRA Learns Less and Forgets Less ■ 논문에서는 programming과 mathematics domain에서 LoRA와 full finetuning의 성능을 비교한다. instruction finetuning (≈100K prompt-response pairs)과 continued pretraining (≈20B unstructured tokens)이라는 두 가지 데이터 학습 설정을 다룬다. ■ 실험 결과, 일반적으로 사용되는 낮은 rank 설정에서는 LoRA가 full finetuning보다 상당히 낮은 성능을 보인다. 그럼에도 불구하고 LoRA는 target domain 밖의 tasks에서 base model의 기존 성능을 더 잘 유지한다. ■ LoRA는 weight update를 low-rank로 제한하기 때문에, 이 논문의 cod..
Hyper-Connections ■ 논문에서는 residual connection의 대안으로 사용할 수 있는 간단하면서도 효과적인 방법인 hyper-connections를 제안한다. ■ residual connection의 여러 variants에서 관찰되는 일반적인 문제점들, 예를 들어 gradient vanishing과 representation collapse 사이에서 시소(seesaw) 효과를 해결하는 것을 목표로 한다. ■ 위의 그림과 같은 hyper-connections는 네트워크가 서로 다른 깊이에 위치한 feature들 사이의 connection strength를 조정할 수 있게 해준다. [2409.19606] Hyper-Connections Hyper-ConnectionsWe present hyper-connectio..
Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning ■ golden-standard Chain-of-Thought (CoT) rationales로 LLM을 fine-tuning하거나, 이를 few-shot prompting에서 correct examples로 사용하는 것이 LLM에 도움이 된다는 것을 보여주는 연구들이 등장했다. ■ 인간은 correct examples을 모방하여 학습할 수도 있지만, 자신의 mistakes로부터 학습하는 것 역시 human cognition의 또 다른 중요한 측면이다. ■ 논문에서는 "LLM은 자신의 mistakes로부터 학습했을 때, reasoning 능력 측면에서 그 mistakes로부터 이점을 얻을 수 있는가?"에 초점을 둔다. 그리고 이 문제를 prompting과 model-tuning이라는 두 관점에서 확인한다...
[DeepSeek-V3.2] Pushing the Frontier of Open Large Language Models ■ DeepSeek-V3.2의 key technical breakthroughs는 다음과 같다: (1) efficient attention mechanism인 DeepSeek Sparse Attention (DSA) (2) Scalable Reinforcement Learning Framework (3) Large-Scale Agentic Task Synthesis Pipeline [2512.02556] DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models DeepSeek-V3.2: Pushing the Frontier of Open Large Language ModelsWe introduce DeepSeek-V3.2, a model tha..
[Self-Error-Instruct] Generalizing from Errors for LLMs Mathematical Reasoning ■ LLM은 mathematical reasoning에서는 여전히 많은 bad cases에서 어려움을 겪는다. 기존의 approaches은 bad case 하나하나에서만 training data를 확장하여 생성하기 때문에, 이러한 bad case들에 내재된 더 넓은 error pattern을 충분히 generalize하지 못한다. ■ 논문에서는 더 generalizable하면서도 targeted된 training data를 합성하기 위한 framework인 Self-Error-Instruct (SEI)를 제안한다. [2505.22591] Self-Error-Instruct: Generalizing from Errors for LLMs Mathematical Reasoning Self-Error-Ins..
[SDFT] Self-Distillation Bridges Distribution Gap in Language Model Fine-Tuning ■ 논문에서는 language model을 downstream task에 fine-tuning하는 과정에서 발생하는 catastrophic forgetting과, fine-tuning 과정에서 발생하는 distribution shift가 general task capability뿐 아니라 model의 safety alignment와 helpfulness까지 저하시킬 수 있음을 보여준다. ■ 저자들은 task dataset과 LLM 사이의 distribution gap이 이러한 문제를 일으키는 주된 근본 원인이라고 가정하며, 이 문제를 해결하기 위한 방법으로 Self-Distillation Fine-Tuning (SDFT)을 제안한다. ■ SDFT는 model 자신이 생성한 distilled datase..
[Self-Instruct] Aligning Language Models with Self-Generated Instructions ■ large "instruction-tuned" language models, 즉 instructions에 응답하도록 fine-tuning된 language models은 새로운 tasks에 zero-shot으로 generalize하는 뛰어난 능력을 보여 왔다. ■ 그럼에도 이러한 model들은 human-written instruction data에 크게 의존하며, 이러한 data는 quantity, diversity, creativity가 제한되어 있는 경우가 있기 때문에 tuning된 model의 generality를 저해한다. ■ 논문에서는 pretrained language model 자신의 generation을 이용해 bootstrapping함으로써 instruction-following c..