Human preference alignment is a crucial step in the training process of most visual generative models, be it image-, video-, or, more recently, even world models [1,2,3,4]. Without it, outputs are typically less pleasing, less natural, and less coherent [3, 5, 6].

Existing approaches for aligning to human preferences typically make one of two compromises. Methods that directly optimize on human preferences, such as Diffusion-DPO [5], generally collect them offline, before optimization begins, meaning that the feedback distribution remains fixed as the policy changes. Online methods, e.g., utilising Group Relative Policy Optimization, that get new feedback on freshly generated outputs during training, such as Flow-GRPO [7], have been proposed as a solution to this. These have been shown to outperform their offline counterparts, including Flow-DPO and online variants of DPO [7], consistent with similar findings for LLM preference fine-tuning [8]. However, these methods generally rely on learned reward models, such as PickScore [9] or ImageReward [6], to provide the feedback signal. These proxies can diverge from human judgement, especially when the output distribution shifts away from the reward model's training data, and they may be exploited during optimization (reward hacking), for example by producing artifact-prone images that still score highly [10] or by collapsing visual diversity [7].
The obvious solution to the above-mentioned problems would simply be to collect actual human feedback during training and use that as the training signal. However, speed and latency become a huge factor, since you need to collect a lot of preference data in a very short amount of time to keep up with the training loop and avoid extensive (and expensive) idle compute time. At Rapidata, we have built a setup that supports exactly that, enabling a completely new paradigm for post-training of visual generative models. Below we broadly describe how this approach can be, and has been, used for training large-scale visual generative models.
The training algorithm
In this rundown, we use Flow-GRPO [7] as the basis and illustrate how real human feedback fits into the update step of this algorithm. However, the approach is easily adapted to other formulations, such as forward-process methods like AdvantageFlow [11] and Advantage Weighted Matching [12], or methods that work directly with pairwise preferences like Pref-GRPO [13]. (Get in touch if you are interested or need assistance on that part.) For the sake of brevity, we only give a brief, but sufficient overview of the Flow-GRPO algorithm and objective. For the details, we refer to the Flow-GRPO paper [7]. At a high level, the update step works as follows:
- Generate images in groups of G outputs from the same prompt.
- Grade the images to obtain a "preference score" for each image. This is typically done with a reward model, optionally combined with other automated metrics.
- Calculate the advantages for each image within its group. This is essentially a normalized score of how much higher or lower each image scored relative to the mean of the group.
- Update the model weights to make "advantageous" outputs more likely.
To use real human feedback to grade the outputs, Rapidata's API slots in seamlessly at step 2 to replace the reward model. Everything else can be kept the same. The following section explains the Rapidata system and how to set it up for real-time human feedback.
The Rapidata API and Flows
Rapidata provides programmatic access to a massive crowd of annotators for short-form tasks, which is perfect for validating visual generative outputs and collecting preference data. The system generally works on-demand through an API; the feature built specifically for real-time feedback for online RLHF is called Flows. At a high level, Flows enable an easy and efficient way of ranking groups of images against each other by breaking the ranking into pairwise comparisons, which are combined into final scores through the Bradley–Terry model [14]. As a result, each image in a group gets an associated Elo-style score, which can be used directly to calculate the advantage values. The high-level structure is illustrated in the figure below. Note that the potentially hundreds of different groups are rated asynchronously in parallel, which allows minimizing GPU idle time while the ratings are being collected. That is, rating can start before all images of a batch have been generated.

Below is the minimum code to set up and run Flows. For more details, we refer to our documentation.
The first step is to create the Flow itself. The same Flow should be reused for a training run, so this only needs to be done once, before starting the actual training loop:
from rapidata import RapidataClient
client = RapidataClient()
flow = client.flow.create_ranking_flow(
name="Image Quality Ranking",
instruction="Which image looks better?",
)This sets up the Flow entity, to which image groups can be sent using the default settings. The most important configuration here is the instruction on which the images are rated. Additional, more advanced configuration options are available (for example, the target, minimum, and maximum number of responses per group), which may need to be tweaked for best performance on a particular use case. We recommend getting in touch with us for assistance with this.
At this point, image groups can be submitted for rating:
flow_item = flow.create_new_flow_batch(
datapoints=[
"https://example.com/image_a.jpg",
"https://example.com/image_b.jpg",
"https://example.com/image_c.jpg",
],
# context="A red fox sitting in a snowy forest at sunrise", # OPTIONAL; the generation prompt
time_to_live=180, # seconds; returns whatever has been collected by then
)The images can be supplied either as URLs or local paths. The generation prompt can be optionally passed as context if you want the annotators to judge prompt adherence. Rating starts immediately once the images have been submitted, and the results can be retrieved with:
results = flow_item.get_results()
scores = results.datapoints # {"https://example.com/image_a.jpg": 1243, ...}get_results() blocks until the group is complete (or its time_to_live has expired), and datapoints maps each image to its score. Turning these into group-relative advantages is then the same normalization Flow-GRPO applies to reward-model outputs:
import numpy as np
image_urls = [...] # the group, in the order the policy generated it
rewards = np.array([scores[url] for url in image_urls], dtype=np.float32)
advantages = (rewards - rewards.mean()) / (rewards.std() + 1e-8)Because the scores are normalized within each group, the absolute Elo scale does not matter. The submit and retrieve calls are made repeatedly as part of the training loop, and that is basically all that is needed to replace a reward model with real, online human feedback. Below are a few notes on more advanced implementation details that may be worth considering for larger training runs.
Advanced Implementation Details
- LoRAs: Updating the entire model during alignment post-training is uncommon and often undesirable. The much more efficient approach is to train a LoRA [15, 16]. There is plenty of good literature on post-training with LoRAs, so we will not dive further into it here.
- Optimizing idle time: While the system is extremely fast and can consistently return ratings on sub-minute deadlines, one might still want to eliminate idle compute time entirely. This can be done by shifting human evaluation back one step, so that feedback on one batch is collected while the next batch of image groups is already being generated.
- How much throughput is needed? The human feedback bandwidth you need is proportional to your compute: the more images you generate, the more feedback you need. Here is a quick back-of-the-napkin calculation based on the Flow-GRPO setup [7]. They train on 24 A800 GPUs, and in each round they generate 48 groups of 24 images. That is, 48 Flow items (image groups) per training step. From experience, we find that for groups of 24 images, 100–150 responses per group is reasonable. At a deadline of 3 minutes, which is a reasonable estimate when accounting for the necessary forward and backward passes, that results in a feedback bandwidth of 2,400 responses per minute (48 × 150 / 3). This is easily achievable with the Rapidata system. One can of course also stagger this, e.g., by having 3 batches of 16 groups each finish in one minute.
- Distributed training: When hundreds or thousands of GPU workers query the Flow, having each one authenticate separately can trigger a burst of simultaneous token refreshes and get you rate-limited. Instead, authenticate once in a coordinator process (e.g. rank 0) and share the token. The simplest option is a token file on shared storage: the coordinator calls
maintain_token_file(path), and each worker usesRapidataClient(token_file=path). Treat this file like any other credential: restrict its permissions to your training job and keep it out of version control and logs. For custom setups, see the distributed training guide.
Getting started
Online RLHF with real human feedback removes the need for a proxy reward model and keeps the feedback signal aligned with what people actually prefer, even as the model's output distribution shifts during training. With Flows, it fits into an existing GRPO-style training loop with just a few lines of code. If you are interested in running training like this, or want help adapting it to your algorithm, get in touch at info@rapidata.ai. If you are working on this in an academic context, you may also qualify for our research grants, which provide up to $50,000 in annotation credits per project, with more available for topics close to our own research, such as online RLHF.
References
[1] ByteDance Seed. Seedream 3.0 Technical Report. arXiv:2504.11346, 2025. https://arxiv.org/abs/2504.11346
[2] Qwen Team. Qwen-Image-2.0-RL Technical Report. arXiv:2606.27608, 2026. https://arxiv.org/abs/2606.27608
[3] ByteDance Seed. Seedance 1.0: Exploring the Boundaries of Video Generation Models. arXiv:2506.09113, 2025. https://arxiv.org/abs/2506.09113
[4] Liu, J., Liu, G., Liang, J., Yuan, Z., Liu, X., Zheng, M., et al. Improving Video Generation with Human Feedback. arXiv:2501.13918, 2025. https://arxiv.org/abs/2501.13918
[5] Wallace, B., Dang, M., Rafailov, R., Zhou, L., Lou, A., Purushwalkam, S., Ermon, S., Xiong, C., Joty, S., & Naik, N. Diffusion Model Alignment Using Direct Preference Optimization. CVPR 2024, pp. 8228–8238. https://arxiv.org/abs/2311.12908
[6] Xu, J., Liu, X., Wu, Y., Tong, Y., Li, Q., Ding, M., Tang, J., & Dong, Y. ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation. NeurIPS 2023. https://arxiv.org/abs/2304.05977
[7] Liu, J., Liu, G., Liang, J., Li, Y., Liu, J., Wang, X., Wan, P., Zhang, D., & Ouyang, W. Flow-GRPO: Training Flow Matching Models via Online RL. NeurIPS 2025. arXiv:2505.05470. https://arxiv.org/abs/2505.05470
[8] Tajwar, F., Singh, A., Sharma, A., Rafailov, R., Schneider, J., Xie, T., Ermon, S., Finn, C., & Kumar, A. Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data. ICML 2024. https://arxiv.org/abs/2404.14367
[9] Kirstain, Y., Polyak, A., Singer, U., Matiana, S., Penna, J., & Levy, O. Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image Generation. NeurIPS 2023. https://arxiv.org/abs/2305.01569
[10] Hong, Y., Kao, K.-C., Zhou, H., & Hsieh, C.-J. Understanding Reward Hacking in Text-to-Image Reinforcement Learning. arXiv:2601.03468, 2026. https://arxiv.org/abs/2601.03468
[11] Kveton, B., Rao, A., Mukherjee, S., Singh, K. K., & Lai, V. D. AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models. arXiv:2605.26013, 2026. https://arxiv.org/abs/2605.26013
[12] Xue, S., Ge, C., Zhang, S., Li, Y., & Ma, Z.-M. Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models. arXiv:2509.25050, 2025. https://arxiv.org/abs/2509.25050
[13] Yang, K., Tao, J., Lyu, J., Ge, C., Chen, J., Li, Q., Shen, W., Zhu, X., & Li, X. Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model. CVPR 2024, pp. 8941–8951. https://arxiv.org/abs/2311.13231
[14] Bradley, R. A., & Terry, M. E. Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons. Biometrika, 39(3/4), 324–345, 1952.
[15] Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. LoRA: Low-Rank Adaptation of Large Language Models. ICLR 2022. https://arxiv.org/abs/2106.09685
[16] Fan, Y., Watkins, O., Du, Y., Liu, H., Ryu, M., Boutilier, C., Abbeel, P., Ghavamzadeh, M., Lee, K., & Lee, K. DPOK: Reinforcement Learning for Fine-tuning Text-to-Image Diffusion Models. NeurIPS 2023. https://arxiv.org/abs/2305.16381

