Replies: 8 comments 1 reply
0 replies
|
大佬慢点,我先入个门😭 |
0 replies
|
我认领这篇:Aligning Language Models with Offline Reinforcement Learning from Human Feedback,研究研究,到时候分享一个ppt |
1 reply
|
我认领这篇:DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales |
0 replies
|
我认领openAI这篇吧:《Training language models to follow instructions with human feedback》,到时候一起ppt圆桌分享哇。 |
0 replies
0 replies
|
我认领这篇:Improving alignment of dialogue agents via targeted human judgements |
0 replies
|
考虑支持视觉的RLHF-V吗? |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Discuss the Implementation of RLHF
All reactions