Lifestyle

Google Cloud explains how businesses can customize Gemini with reinforcement learning fine-tuning (RLFT), training on scores instead of model answers

Google Cloud has published a guide to its managed reinforcement learning fine-tuning service. Businesses and developers write a program that scores Gemini's answers, and the service tunes the model toward higher scores. The guide covers when this helps more than other methods and how early adopters used it, as described by Google Cloud.

About 6 min read

Google Cloud explains how businesses can customize Gemini with reinforcement learning fine-tuning (RLFT), training on scores instead of model answers
Image: Mokaair (Original editorial artwork)

What happened

The Google Cloud blog published a guide by Google senior software engineer Jiaqi Pan and senior product manager Kunal Jha on customizing Gemini models with reinforcement learning (RL). RL is a way of training a model by rewarding good results. Google Cloud notes that RL is key to post-training modern large language models, meaning the extra training a model gets after its initial build. It adds that RL requires large training clusters and access to model internals, which external customers cannot get for proprietary models such as Gemini. Google Cloud says it has therefore packaged RL as a managed fine-tuning service. Customers only provide prompts and a reward function, and Google takes care of the infrastructure and model internals.

As Google Cloud describes it, at each training step the service generates several candidate responses and scores them with the user's reward. It then adjusts the model so that high-scoring responses become more likely, while keeping the model close to the original Gemini. The RL details are fully managed. What the user is responsible for, and what most affects the outcome, is the reward itself.

Google Cloud explains how businesses can customize Gemini with reinforcement learning fine-tuning (RLFT), training on scores instead of model answers
Mokaair editorial verification flow · Image: Mokaair (Original editorial artwork)
Read the full description

Sources are collected, independently checked, then reviewed by Jev.

How RLFT differs from SFT

Supervised fine-tuning (SFT) gives the model a set of example answers to imitate. RLFT, as Google Cloud describes it, instead means writing a grading program so the model improves toward higher scores. Google Cloud lists three characteristics of RLFT. First, it learns from the model's own outputs, so it tends to disturb unrelated capabilities less than SFT. Second, it rewards outcomes rather than the path taken to reach them. Third, it amplifies abilities the model already has, turning occasional successes into reliable ones, but it cannot teach the model skills it has never shown.

Comparison based on Google Cloud's guide
AspectSupervised fine-tuning (SFT)Reinforcement learning fine-tuning (RLFT)
Learns fromLabeled example answersA user-defined reward signal
Suited tasksTasks where examples are easy to provideTasks hard to demonstrate but easy to score
Multiple correct answersA single reference answer may wrongly penalize other correct solutionsAny good outcome can score
Impact on unrelated capabilitiesRelatively larger, per Google CloudOften smaller, per Google Cloud
LimitationMay plateau on key metricsCannot teach skills the model has never shown

Google Cloud recommends trying prompting and SFT first. If the base model already succeeds some of the time, RLFT can be used directly. If the success rate is too low, or you already have SFT data, Google Cloud suggests two stages. Start with a short, light round of SFT as a warm start. Then continue with RL from the saved SFT version of the model, known as a checkpoint, using a feature called Continuous Tuning.

Use cases cited by Google Cloud

  • Game NPCs (computer-controlled characters): a Gemini autorater, where one AI model grades another (LLM-as-a-judge), scores role-play and language. Google Cloud says repetitive loops and drifting into the wrong language disappeared.
  • Structured entity extraction, meaning pulling items such as invoice fields out of documents: the reward is rule-based and scores precision (no invented fields) and recall (no missing fields). Google Cloud says field accuracy improved on noisy real-world documents, turning a manual review step into an automated one.
  • Content moderation: a reward running on Cloud Run combines format validation with a deterministic grader, one that always gives the same verdict for the same input. Google Cloud says false positives, meaning content wrongly flagged, dropped sharply, and fewer cases had to be escalated to humans.
  • Code measured by execution: SQL or API calls are run in a secure sandbox and then scored. Google Cloud says this lets non-technical users query proprietary data in natural language.
  • HTML slide generation: slides are rendered and then scored on layout and design. Google Cloud says the output was well-styled decks with cohesive themes and no layout overflow.

Practical impact for general readers and businesses

For everyday users, this kind of technology may make AI tools inside companies more reliable at specific jobs. Examples include extracting data from invoices more accurately, or letting people who cannot code query company data in natural language. For enterprise developers, Google Cloud stresses that reward design is key. A good reward should match what humans prefer and handle badly formatted outputs without crashing. It should also resist "reward hacking", where the model exploits loopholes in the grading instead of actually completing the task. Google Cloud also recommends strictly separating training and evaluation data and starting from default settings. Finally, it advises keeping the checkpoint where the validation reward stops improving, rather than the one from the final step.

Frequently asked questions

What is RLFT?

According to Google Cloud, RLFT is a managed reinforcement learning fine-tuning service. It lets customers tune Gemini with custom reward functions instead of labeled answers, with Google handling the RL infrastructure and model internals.

Do I have to use RLFT to customize Gemini?

No. Google Cloud recommends trying prompting and supervised fine-tuning first. It says RLFT is most valuable when outputs can be graded but good answers are hard to write, when SFT has plateaued, or when a task has many equally valid answers.

Can RLFT teach a model entirely new skills?

According to Google Cloud, no. RLFT amplifies existing abilities, making occasional successes reliable, but it cannot teach skills the model has never shown. When the success rate is too low, you can warm-start with SFT first.

What is reward hacking?

It is when the model finds loopholes in how it is scored and gets high marks without actually completing the task. Google Cloud recommends using multiple judges, penalizing overly long outputs, and favoring verifiable checks over relying on a model's opinion.

Are the reported results trustworthy?

These cases are early-adopter experiences as described by Google Cloud. No specific figures were published and the results have not been independently verified, so they should be treated as vendor claims.

Browse the latest news in this topic

Latest travel guides

Sources

Lifestyle