Reinforcement Learning from Human Feedback: LLM alignment and post-training

Author: Nathan Lambert

Language: English

Genre: Artificial Intelligence