Skip to content
OTFotf
All posts

A safe evaluation plan for Vercel’s free Glyph Cluster model

D
DaveAuthor
8 min read
A safe evaluation plan for Vercel’s free Glyph Cluster model

A free model is still a production decision. Glyph Cluster is available through Vercel AI Gateway during a stealth period, but the launch notice says its provider may use prompts and responses for training and model improvement, and that zero data retention (ZDR) is unavailable. The model page calls it anonymous and says prompts and outputs may be retained by the provider. Those are material constraints for any app that handles private code, customer records, credentials, or unreleased product plans.

The first question is whether you can test it without sending information you cannot afford to disclose. If you cannot isolate public or owned non-sensitive examples, keep it out of the evaluation. If the test passes, decide whether behavior, output format, and measured latency earn a narrow place in your stack. Do not make “free” the production gate.

A team checks sample documents at a review gate that blocks private data before a model evaluation

Start with the data boundary

Vercel’s October 7 announcement says Glyph Cluster is a stealth model for AI Gateway customers on Pro and Enterprise plans with purchased Gateway credits, and that it is free during the stealth period. The same announcement says ZDR is not available and prompts and responses may be used for training and model improvement. The model page adds that Glyph Cluster is anonymous and that its provider may retain prompts and outputs.

There is a source conflict to preserve, not smooth over. The changelog warns that prompts and responses may be used for training. The current AI Gateway model listing also has a “No Training” indicator, but its Stealth provider row does not carry that marking. That is a conflict in the vendor’s own published signals. Treat the explicit changelog warning as the operative constraint for a test, and ask Vercel for clarification before relying on a different interpretation. “Anonymous” also means the provider is not identified on the public model page. Stealth access and its terms can change.

That makes the initial test boundary simple: use only public material or non-sensitive examples that your team owns and is allowed to send to this service. Do not use a private repository, customer data, access tokens, incident reports, unreleased designs, or real support conversations. Do not assume that redaction makes a sensitive prompt safe; snippets can retain context that identifies a customer or exposes implementation details.

If you cannot produce a representative test set within that boundary, stop. A model that cannot be safely evaluated for your workload is not a candidate for that workload yet.

Pick one task where the output can be inspected

Vercel describes Glyph Cluster as a reasoning model for coding and long-context analysis, including planning, synthesis, quantitative reasoning, and comparing material across large inputs. It supports function calling but not structured outputs; input is text only, without image or file input. Start with a text task whose output a human can review against a clear reference.

Good first candidates are tasks such as explaining a public build failure, comparing two owned non-sensitive implementation plans, or suggesting a patch against a small sample repository that contains no secrets. Avoid using it first for automatic customer-facing actions, permission decisions, payments, or any flow that needs a schema-validated response. Function calls do not change the format limitation: the model can call tools you define, but the launch page does not claim support for structured output.

Make the task contract explicit before comparing models:

  • What evidence must the answer cite from the input?
  • What would count as a wrong or incomplete result?
  • Which outputs need a human decision before they can affect a user or system?
  • What is the maximum acceptable latency for this path?
  • Which data classes are excluded from the test?

11 production screens. Login, database, payments — all wired.

The SaaS Dashboard Kit ships everything already connected. Nothing to set up. Live demo at saas.otf-kit.dev.

See the live demo

Run a small, controlled evaluation

Build a compact set of public or owned, non-sensitive examples that resembles the task without reproducing private production records. Include routine cases and known edge cases. Write expected outcomes before running either model. For code review, that could mean a seeded bug with a known fix, an innocuous code sample with no defect, and a case where the right answer is to ask for more context. For long-context comparison, make sure the correct answer depends on facts spread across the text and record which facts must be found.

Keep the test behind a development-only route or a local script. Use a server-side key in an environment variable, never a browser bundle or committed file. Call the Gateway model by its documented identifier, stealth/glyph-cluster, and save only the non-sensitive test input, output, model ID, timestamp, and measurements you need. Do not turn on function tools that can modify systems while evaluating; if a tool call is required to demonstrate a task, use a read-only stub or mock.

For example, with the AI SDK’s generateText API shown in Vercel’s model documentation:

import { generateText } from 'ai';

const result = await generateText({
  model: 'stealth/glyph-cluster',
  prompt: `Review this non-sensitive sample and list:
1. The likely failure.
2. The evidence in the input.
3. A minimal proposed fix.
If the evidence is insufficient, say what is missing.

${sample}`,
});

console.log(result.text);

This gives you text to inspect; it does not provide a schema-validated object. Validate the response in your own code before using it, and keep any subsequent action behind your existing human or deterministic checks. If your product requires machine-readable fields, this model’s documented lack of structured outputs is a reason to choose another route or add a separate validation step—not to parse free-form text optimistically.

Measure quality and latency, not just a good example

Run the same cases against your current model and Glyph Cluster. Score each output against the expected result, not by whether it sounds confident. For coding, check whether the diagnosis matches the seeded issue, whether the proposed change is minimal, and whether it introduces a new failure when applied to a disposable sample. For synthesis, check every required fact and mark unsupported claims. Record a few operational measures for each call: success or error, time to first token if you stream, total response time, and output length. Vercel’s model page currently shows a live provider row and latency values, but those are a snapshot, not a promise for your traffic. Measure the route you will actually use, with the same request shape and network path. Treat the stealth provider’s identity and availability as mutable; do not pin a design to assumptions that the public page does not make.

MeasureKeep if…Stop or narrow if…
CorrectnessIt matches the written expected result on routine and edge casesIt misses required evidence or invents a fix
Review effortA reviewer can verify the output quicklyFree-form answers take longer to validate than the current route
LatencyThe measured path fits your product’s limitTail cases exceed the user-facing budget
Data boundaryEvery test input is approved, public, or non-sensitive and ownedA representative case needs private or customer data
Output contractYour code can safely validate the text resultThe feature requires structured output that is not supported

This scorecard is a decision tool, not evidence that Glyph Cluster outperforms another model; only your controlled test can establish fit.

Byte and Luna compare two text-only evaluation results using a scorecard and stopwatch before a gated rollout

Keep any rollout narrow and reversible

If the evaluation passes, start with a low-risk internal or opt-in path that still uses non-sensitive inputs. Keep the existing model as the default. Route only the named task to Glyph Cluster, log its model ID and outcome, and retain a quick way to disable the route. Do not send sensitive data merely because the task has passed on public examples. A later decision to use a different data class needs a separate review of the provider’s identity, retention, training terms, and your own obligations.

Recheck the model page and terms before expanding access. The model is labeled stealth and anonymous, the provider may retain inputs and outputs, ZDR is unavailable according to the launch notice, and the model’s terms can change. Free access is temporary, so also decide what happens when the stealth period ends: remove the route, replace it, or obtain a confirmed price and service commitment before continuing.

The go/no-go rule

Try Glyph Cluster only when the task is text-based, the test can use public or owned non-sensitive examples, and a reviewer can verify its output. Put the training and retention warning, missing ZDR, anonymous provider, unsupported structured outputs, and changing stealth terms ahead of any quality claim. Measure against your current route on a fixed set, then keep the experiment narrow and reversible.

If those conditions do not fit the task, do not route it there because the current price is zero. Wait for clarified terms or choose a model whose data handling and output contract match the workload. For a model-selection workflow that tests the API identifier rather than trusting a label, see our API pinning guide.

Sources

ai-toolsarchitecturevercel
OTF SaaS Dashboard Kit

Ship the product, not the setup.

  • 11 production screens — auth, billing, team, analytics, settings
  • Real database, payments, and login — all wired on day 1
  • AI configs pre-tuned so your agent extends instead of regenerates
Need more than components?

Full-stack kits.
Pay once, own the code.

Auth, database, and payments already connected — so you ship product, not setup. Or take the delivered kits in the Bundle.

Everything Bundle — $149See full pricing

Get the free AI configs pack

Pre-tuned AI configs for Cursor, Claude, and Lovable — drop them in and your AI tool instantly understands your project.

No spam. Unsubscribe any time.

Prefer the free SDK? Star it on GitHub →