Lifestyle
Google Cloud explains how businesses can customize Gemini with reinforcement learning fine-tuning (RLFT), training on scores instead of model answers
Google Cloud has published a guide to its managed reinforcement learning fine-tuning service. Businesses and developers write a program that scores Gemini's answers, and the service tunes the model toward higher scores. The guide covers when this helps more than other methods and how early adopters used it, as described by Google Cloud.
About 6 min read

What happened
The Google Cloud blog published a guide by Google senior software engineer Jiaqi Pan and senior product manager Kunal Jha on customizing Gemini models with reinforcement learning (RL). RL is a way of training a model by rewarding good results. Google Cloud notes that RL is key to post-training modern large language models, meaning the extra training a model gets after its initial build. It adds that RL requires large training clusters and access to model internals, which external customers cannot get for proprietary models such as Gemini. Google Cloud says it has therefore packaged RL as a managed fine-tuning service. Customers only provide prompts and a reward function, and Google takes care of the infrastructure and model internals.
As Google Cloud describes it, at each training step the service generates several candidate responses and scores them with the user's reward. It then adjusts the model so that high-scoring responses become more likely, while keeping the model close to the original Gemini. The RL details are fully managed. What the user is responsible for, and what most affects the outcome, is the reward itself.
Read the full description
Sources are collected, independently checked, then reviewed by Jev.
How RLFT differs from SFT
Supervised fine-tuning (SFT) gives the model a set of example answers to imitate. RLFT, as Google Cloud describes it, instead means writing a grading program so the model improves toward higher scores. Google Cloud lists three characteristics of RLFT. First, it learns from the model's own outputs, so it tends to disturb unrelated capabilities less than SFT. Second, it rewards outcomes rather than the path taken to reach them. Third, it amplifies abilities the model already has, turning occasional successes into reliable ones, but it cannot teach the model skills it has never shown.
| Aspect | Supervised fine-tuning (SFT) | Reinforcement learning fine-tuning (RLFT) |
|---|---|---|
| Learns from | Labeled example answers | A user-defined reward signal |
| Suited tasks | Tasks where examples are easy to provide | Tasks hard to demonstrate but easy to score |
| Multiple correct answers | A single reference answer may wrongly penalize other correct solutions | Any good outcome can score |
| Impact on unrelated capabilities | Relatively larger, per Google Cloud | Often smaller, per Google Cloud |
| Limitation | May plateau on key metrics | Cannot teach skills the model has never shown |
Google Cloud recommends trying prompting and SFT first. If the base model already succeeds some of the time, RLFT can be used directly. If the success rate is too low, or you already have SFT data, Google Cloud suggests two stages. Start with a short, light round of SFT as a warm start. Then continue with RL from the saved SFT version of the model, known as a checkpoint, using a feature called Continuous Tuning.
Use cases cited by Google Cloud
- Game NPCs (computer-controlled characters): a Gemini autorater, where one AI model grades another (LLM-as-a-judge), scores role-play and language. Google Cloud says repetitive loops and drifting into the wrong language disappeared.
- Structured entity extraction, meaning pulling items such as invoice fields out of documents: the reward is rule-based and scores precision (no invented fields) and recall (no missing fields). Google Cloud says field accuracy improved on noisy real-world documents, turning a manual review step into an automated one.
- Content moderation: a reward running on Cloud Run combines format validation with a deterministic grader, one that always gives the same verdict for the same input. Google Cloud says false positives, meaning content wrongly flagged, dropped sharply, and fewer cases had to be escalated to humans.
- Code measured by execution: SQL or API calls are run in a secure sandbox and then scored. Google Cloud says this lets non-technical users query proprietary data in natural language.
- HTML slide generation: slides are rendered and then scored on layout and design. Google Cloud says the output was well-styled decks with cohesive themes and no layout overflow.
Practical impact for general readers and businesses
For everyday users, this kind of technology may make AI tools inside companies more reliable at specific jobs. Examples include extracting data from invoices more accurately, or letting people who cannot code query company data in natural language. For enterprise developers, Google Cloud stresses that reward design is key. A good reward should match what humans prefer and handle badly formatted outputs without crashing. It should also resist "reward hacking", where the model exploits loopholes in the grading instead of actually completing the task. Google Cloud also recommends strictly separating training and evaluation data and starting from default settings. Finally, it advises keeping the checkpoint where the validation reward stops improving, rather than the one from the final step.
Frequently asked questions
What is RLFT?
According to Google Cloud, RLFT is a managed reinforcement learning fine-tuning service. It lets customers tune Gemini with custom reward functions instead of labeled answers, with Google handling the RL infrastructure and model internals.
Do I have to use RLFT to customize Gemini?
No. Google Cloud recommends trying prompting and supervised fine-tuning first. It says RLFT is most valuable when outputs can be graded but good answers are hard to write, when SFT has plateaued, or when a task has many equally valid answers.
Can RLFT teach a model entirely new skills?
According to Google Cloud, no. RLFT amplifies existing abilities, making occasional successes reliable, but it cannot teach skills the model has never shown. When the success rate is too low, you can warm-start with SFT first.
What is reward hacking?
It is when the model finds loopholes in how it is scored and gets high marks without actually completing the task. Google Cloud recommends using multiple judges, penalizing overly long outputs, and favoring verifiable checks over relying on a model's opinion.
Are the reported results trustworthy?
These cases are early-adopter experiences as described by Google Cloud. No specific figures were published and the results have not been independently verified, so they should be treated as vendor claims.
Browse the latest news in this topic
Lifestyle
NVIDIA launches DGX Spark 64GB: on sale October 23 from $4,999, two units can be linked into 128GB
On October 2, 2026, NVIDIA announced a more affordable 64GB memory version of its DGX Spark personal AI computer, available from October 23 through six makers including Acer and ASUS. It is aimed mainly at developers and researchers who want to run AI models on their own machines. Below we summarize the specs NVIDIA published, its claims about linking two units, and what it means for general readers.
Lifestyle
Google Cloud Launches Spanner Queues: Putting Message Queues Inside Database Transactions to Make AI Agents More Reliable
Google Cloud has announced the general availability of Spanner queues, which make message creation part of a database transaction. The aim is to stop AI agents' "state" and "actions" from falling out of sync. This article covers Google Cloud's claims, the main features, and what it means for general readers.
Lifestyle
GPT-6.1 Sol Launches: New Sol Version in the API, Codex and ChatGPT Work, Not in Chat
OpenAI launched GPT-6.1 Sol on September 29, 2026, with the API name gpt-6.1-sol. The launch rollout covers Codex and ChatGPT Work on Plus, Pro, Business, Enterprise and Edu (Enterprise and Edu need an administrator to enable it); Free and Go are not included at launch, and it is not in Chat (checked September 2026).
Lifestyle
Claude Sonnet 5.5 Launches: Same List Price as Sonnet 5, Available in the API, on Cloud Platforms and in Claude.ai
Anthropic launched Claude Sonnet 5.5 on September 28, 2026. API list prices are the same as Sonnet 5 ($2 per million input tokens, $10 per million output tokens). It is available in Claude.ai, the API and several cloud platforms, and higher-risk cybersecurity requests fall back to Sonnet 5 (checked September 2026).
Latest travel guides

GuideTokyo
Where to Stay in Tokyo: Comparing Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza, Plus Airport Access, Accommodation Tax, and Luggage Delivery
Where should you stay in Tokyo? Compare Shinjuku, Ueno, Tokyo Station, Shibuya, Asakusa, Ikebukuro, and Ginza by the same criteria: access from Narita and Haneda, transit routes, nearby attractions, neighborhood character, and who each area suits. Includes a comparison table, a Yamanote Line diagram, Tokyo’s accommodation tax as verified in 2026/9 (changing to 3% in 2027/4), and Airport TA-Q-BIN luggage shipping rules.
- Budget
- Hotels

GuideTokyo
How to Choose Tokyo Transit Passes: Are Suica, Welcome Suica, the Tokyo Subway Ticket, and the JR Pass Worth It?
On a first Tokyo trip, start with an IC card and pay per ride (Welcome Suica has no deposit and is valid for 28 days). If you take four or more subway rides in a day, add a 72-hour Tokyo Subway Ticket for 2,000 yen; a JR Pass is never worthwhile if you stay in Tokyo and do not go to Kansai. See what TOURIST PASMO, Suica on iPhone, and the Tokyo Metro day pass do and do not cover, with a decision chart. Prices verified in September 2026.
- Transport
- Budget

GuideTokyo
Tokyo Disneyland and DisneySea Guide: Ticket Prices, Fantasy Springs, Disney Premier Access (DPA), Standby Pass, and Which Park to Choose for Your First Visit
Tokyo Disney one-day Passport prices vary: most weekdays in 9/2026 cost ¥9,900 and weekends ¥10,900. At 14:00 daily, tickets go on sale for the same date two months later. Free Priority Pass is no longer on the official service list; only paid Disney Premier Access (¥1,000–3,500 per person per use) shortens waits. Covers hours, the 25th anniversary, Standby Pass, Entry Request, Fantasy Springs access and first-visit park choice; checked on the official site in 9/2026.
- Itineraries
- Family