7 entries · 4 contributors45 minutes of reading, totalLast update · Sep 30, 2026

Research

Long reads, experiments and insights. Written by the people who build Rapidata — for the people who train models.

№ 01the lead

Live Human Feedback in the Training Loop: Aligning Diffusion Models with Real People

Existing approaches for aligning to human preferences typically make one of two compromises. Methods that directly optimize on human preferences, such as Diffusion-DPO [5],…

Mads Kuhlmann-JoergensenSep 30, 20268 min
Read the lead
// newsletter · monthly

One email a month.

What we’re learning from human feedback at scale.

~2,800 ML practitioners · no spam · unsubscribe in one click
The Index4 more entries
№ 045 min

Landing the Lunar Lander with Human Feedback

Reinforcement Learning from Human Feedback is most commonly associated with the final training stages of large language models like ChatGPT or Claude. It’s what helps these models…

Daniil PyatkoMay 12, 20255 min
№ 056 min

Beyond Image Preferences - Rich Human Feedback

TL;DR: We collected 1.5 million annotations from >150 thousand individual humans using Rapidata via the Python API to build a dataset of detailed human feedback for text-to-image…

Mads Kuhlmann-JoergensenJan 10, 20256 min
№ 073 min

Object Detection in Computer Vision

Object detection in computer vision is a fascinating blend of mathematics, algorithms, and machine learning that allows computers to identify and locate objects within an image or…

Marian KannwischerApr 30, 20243 min