Every new label in a classification system carries a cost. Someone collects examples, annotates them, and retrains the model before the system recognizes that label at all. For a support queue that gains a ticket category with every product launch, or a moderation team facing a manipulation tactic that did not exist last month, that loop runs slower than the problem it is supposed to solve. Zero-shot classification removes the loop. A model assigns input to a class it has never seen a labeled example of, working only from a description of what that class means.

This guide explains how zero-shot classification works, the two model families teams deploy for it, how it compares to supervised learning, and where it breaks down. It also covers the step most ML tutorials skip. Embedding-based zero-shot classification is a vector similarity search, and you can run that search in SQL, next to your application data, with 티DB vector search.

What Is Zero-Shot Classification?

Zero-shot classification is a machine learning technique that assigns input to a category the model was never trained on, using a natural-language description of the category instead of labeled examples. A supervised classifier recognizes only the labels in its training data. A zero-shot classifier scores the input against a description of each candidate label and returns the closest match.

The “zero” counts labeled examples per new class. Few-shot learning needs a handful. Supervised classification needs hundreds or thousands. Zero-shot classification needs none, because the model draws on what it learned while pretraining on large, general datasets. A label you define five minutes before a request arrives is usable on that request.

How Zero-Shot Classification Works

Zero-shot classification works by scoring how well an input matches a description of each candidate label, using knowledge the model picked up during pretraining. Two model families do this in different ways, and the difference decides how you deploy them. Natural language inference (NLI) models score each input and label pair directly. Embedding models convert inputs and labels into vectors and compare the vectors.

NLI models such as BART-MNLI

Natural language inference is the task of deciding whether one sentence (the premise) entails, contradicts, or is neutral toward another (the hypothesis). The facebook/bart-large-mnli model is BART fine-tuned on the Multi-Genre Natural Language Inference (MultiNLI) corpus, and it is the default model behind the Hugging Face zero-shot-classification pipeline. To classify, the model treats the input as the premise and turns each candidate label into a hypothesis such as “This example is about billing.” The entailment probability becomes that label’s score.

from transformers import pipeline

classifier = pipeline("zero-shot-classification", model="facebook/bart-large-mnli")

result = classifier(
    "The checkout page times out every time I apply a discount code.",
    candidate_labels=["billing", "bug report", "feature request", "account access"],
)

print(result["labels"][0], round(result["scores"][0], 3))

The pipeline returns every candidate label sorted by score. The cost model matters in production. The model runs one forward pass per input and label pair, so 20 candidate labels means 20 passes for every input, and none of that work carries over to the next input. There is no label representation to precompute or store.

Embedding models such as CLIP

An embedding is a fixed-length list of numbers that places text or an image in a vector space where similar meanings sit close together. CLIP, which OpenAI trained on 400 million image and text pairs, encodes images and text into the same space. To classify an image, you embed a prompt for each label, such as “a photo of a golden retriever,” embed the image, and pick the label whose embedding sits closest by cosine similarity. Text-only embedding models, such as the sentence-transformers family, apply the same pattern to documents, tickets, and messages.

The cost model is the opposite of NLI. You embed each label description once and reuse it. Every new input costs one encoding pass plus a nearest-neighbor search across the stored label embeddings, and adding a label adds a row, not a forward pass. That search is the operation a vector database exists to run, which is where TiDB enters the workflow later in this guide.

The table below summarizes how the two families differ in production.

NLI models (BART-MNLI) Embedding models (CLIP, sentence-transformers)
How a label is scored Entailment probability for each input and label pair Cosine similarity between input and label vectors
Work per input with N labels N model forward passes One forward pass plus one vector search
Reusable label representation None Label embeddings, computed once and stored
Adding a label Add a string to the candidate list Embed the description and insert one row
Where it fits Small, fixed label sets at modest volume Large or growing label sets, high volume, text and images

What decides zero-shot classification accuracy

Across both families, three components decide how accurate a zero-shot classifier is in practice.

  • The pretrained model. The model can only match concepts it absorbed during pretraining. A general-purpose model knows “refund request” and struggles with internal jargon or narrow clinical terms.
  • The label descriptions. “Billing” is a weak label. “Questions about invoices, charges, refunds, or payment methods” gives the model far more to match against. Rewriting descriptions is the cheapest accuracy fix available.
  • The scoring threshold. Every input gets a best match, including inputs that belong to no class. Set a minimum score or maximum distance, tuned on a few hundred real examples, and route anything outside it to a fallback.

How Zero-Shot Classification Compares to Supervised Learning

Supervised classification needs labeled examples for every class it recognizes. Zero-shot classification needs only a description. When enough labeled data exists for a class, a supervised model trained on it outperforms zero-shot classification on that class, because it learns the class boundaries from examples instead of inferring them from a sentence.

Supervised classification Zero-shot classification
Labeled data per class Hundreds to thousands of examples None
Adding a class Collect data, annotate, retrain, redeploy Write a description
Accuracy on well-represented classes Higher Lower
Classes defined after training Not supported without retraining Supported
Main cost Annotation and retraining Inference, plus careful label writing

The decision comes down to volume and stability. A high-volume, stable category justifies the upfront cost of labeling examples, and a supervised model trained on them will win. A low-volume, short-lived, or unpredictable category, such as a new slang term in content moderation or a fraud pattern that appears and disappears within weeks, favors zero-shot classification because a supervised model would never be ready in time.

The two approaches also combine. Teams run zero-shot classification on a new category from day one, send low-confidence results to human review, and use the reviewed results as training data. Once a category has enough labeled examples, a supervised model takes it over.

How to Store and Query Zero-Shot Classification Embeddings with TiDB

Embedding-based zero-shot classification is a nearest-neighbor search. Given an input embedding, find the class-label embedding closest to it. TiDB runs that search in standard SQL with a VECTOR column type, a vector index, and the VEC_COSINE_DISTANCE function, so label embeddings live in the same database as the tickets, products, or transactions you are classifying.

This pattern applies to embedding models such as CLIP and sentence-transformers. NLI models such as BART-MNLI produce no reusable label embeddings, so there is nothing to store. Run them through the pipeline shown earlier.

TiDB vector search is available on TiDB Self-Managed, TiDB Cloud Starter, TiDB Cloud Essential, and TiDB Cloud Dedicated. TiDB Self-Managed and TiDB Cloud Dedicated clusters need TiDB v8.4.0 or later, and v8.5.0 or later is recommended.

Create a table for class-label embeddings

CREATE TABLE class_labels (
  id INT PRIMARY KEY AUTO_INCREMENT,
  label VARCHAR(64) NOT NULL,
  description TEXT NOT NULL,
  model_name VARCHAR(128) NOT NULL,
  embedding VECTOR(384) NOT NULL,
  VECTOR INDEX idx_label_embedding ((VEC_COSINE_DISTANCE(embedding)))
);

The dimension in VECTOR(384) must match your embedding model’s output. The all-MiniLM-L6-v2 sentence-transformers model used below produces 384 dimensions, and CLIP ViT-B/32 produces 512. The vector search index uses HNSW (Hierarchical Navigable Small World), an approximate nearest-neighbor algorithm, built on cosine distance, the same metric the classifier uses. When you define a vector index at table creation, TiDB creates a TiFlash replica for the table automatically, so a TiDB Self-Managed cluster needs at least one TiFlash node.

The model_name column records which model produced each embedding. Embeddings from different models do not share a vector space, and this column tells you which rows to re-embed when you switch models.

Embed label descriptions and classify new input

The script below embeds four label descriptions, stores them, and classifies a support ticket against them. Copy the connection parameters from your TiDB Cloud console.

import pymysql
from sentence_transformers import SentenceTransformer

MODEL_NAME = "sentence-transformers/all-MiniLM-L6-v2"  # 384 dimensions
model = SentenceTransformer(MODEL_NAME)

def to_vector(text):
    # TiDB accepts vectors as strings in the form "[0.1, 0.2, ...]"
    return str(model.encode(text).tolist())

labels = {
    "billing": "Questions about invoices, charges, refunds, or payment methods",
    "bug_report": "Something in the product is broken, slow, or returns an error",
    "feature_request": "A request for functionality the product does not have yet",
    "account_access": "Problems signing in, resetting a password, or managing users",
}

conn = pymysql.connect(
    host="<your_host>",
    port=4000,
    user="<your_user>",
    password="<your_password>",
    database="test",
    ssl_verify_cert=True,
    ssl_verify_identity=True,
)

with conn.cursor() as cur:
    for label, description in labels.items():
        cur.execute(
            "INSERT INTO class_labels (label, description, model_name, embedding) "
            "VALUES (%s, %s, %s, %s)",
            (label, description, MODEL_NAME, to_vector(description)),
        )
    conn.commit()

    ticket = "The checkout page times out every time I apply a discount code."
    query_vector = to_vector(ticket)
    cur.execute(
        "SELECT label, VEC_COSINE_DISTANCE(embedding, %s) AS distance "
        "FROM class_labels "
        "ORDER BY VEC_COSINE_DISTANCE(embedding, %s) "
        "LIMIT 3",
        (query_vector, query_vector),
    )
    for label, distance in cur.fetchall():
        print(f"{label}: {distance:.3f}")

The query returns the three closest labels with their distances. Cosine distance runs from 0 (same direction) to 2 (opposite direction), so smaller means closer. The top row is the classification, and the gap between the first and second rows is a practical confidence signal. A ticket where two labels sit at nearly the same distance deserves a human look. For connection options and loading patterns, see the guide to TiDB’s vector search SQL syntax.

Make sure the classification query uses the vector index

TiDB uses the vector index for nearest-neighbor queries written as ORDER BY a distance function followed by LIMIT, where the distance function matches the one the index was built on and the sort is ascending. The query above qualifies. Ordering by VEC_L2_DISTANCE on a cosine index, or sorting in descending order, skips the index, and SHOW WARNINGS reports the reason. Run EXPLAIN on your classification query before it goes to production.

Join classification results against application data

Storing results in the same database turns classification output into something you can query. A results table keyed to your application data lets you count, filter, and audit classifications with ordinary SQL.

CREATE TABLE ticket_classifications (
  ticket_id BIGINT PRIMARY KEY,
  label VARCHAR(64) NOT NULL,
  distance DOUBLE NOT NULL,
  model_name VARCHAR(128) NOT NULL,
  classified_at DATETIME NOT NULL DEFAULT CURRENT_TIMESTAMP
);

-- Labels whose matches are drifting away from their descriptions this week
SELECT label,
       COUNT(*) AS tickets,
       ROUND(AVG(distance), 3) AS avg_distance
FROM ticket_classifications
WHERE classified_at >= NOW() - INTERVAL 7 DAY
GROUP BY label
ORDER BY avg_distance DESC;

A rising average distance for one label means incoming tickets match its description less well than they used to. That is the signal to rewrite the description or split the label in two, and you see it without exporting anything to a separate vector store. The embeddings still come from your model, not from TiDB. TiDB stores them, searches them, and keeps them next to the data the classifications describe.

Where Teams Use Zero-Shot Classification

Zero-shot classification earns its place wherever new categories appear faster than a team can label training data for them. Five domains show the pattern clearly.

Natural language processing

Support teams use zero-shot classification to route tickets for categories no historical data covers. A feature launch creates a wave of questions the existing ticket history cannot label, and a zero-shot classifier routes them from day one with nothing more than a description of each new category. Sentiment analysis tools apply the same pattern to a new market or language without collecting fresh labeled reviews, since the underlying model already recognizes sentiment-bearing language.

Image and video recognition

Content moderation and security monitoring systems use zero-shot classification to flag objects or content types outside the original training set. A moderation team responding to a new manipulation tactic can describe it and start flagging it the same day, instead of spending weeks collecting and labeling enough examples to retrain a detector.

Healthcare

Medical imaging researchers apply zero-shot classification to rare conditions that never appear often enough to build a labeled dataset, working from a text description of the finding instead of hundreds of training examples. Rare disease detection is the clearest case, since by definition too few historical cases exist to train a conventional classifier. Clinical use requires the same validation as any other diagnostic model. A zero-shot result is a flag for a specialist, not a diagnosis.

Finance

Fraud detection teams use zero-shot classification to catch patterns that match no previously labeled example. Fraud tactics evolve to evade whatever a supervised model learned to catch, so a technique that works from a description of the tactic complements existing fraud models. It does not replace them.

E-commerce

Marketplaces use zero-shot classification to categorize new products the moment a seller publishes a listing, without waiting for purchase history. For a catalog adding thousands of listings a day, that closes the gap between a product going live and appearing in the right searches and recommendations. The same labels feed inventory reporting for categories the merchandising team added last week.

What Zero-Shot Classification Does Well, and Where It Breaks Down

Zero-shot classification trades accuracy on known classes for flexibility on new ones. Its advantages come from removing labeled data. Its challenges come from relying on descriptions, pretraining data, and similarity scores in place of that data.

Advantages of zero-shot classification

  • No labeled-data bottleneck. A new class needs a description, not an annotated dataset. Teams without the budget to label every category they might need can still classify against all of them.
  • Classes defined at runtime. A model classifies input into a class defined moments earlier. In content moderation or emerging fraud patterns, waiting for a retrained model is not an option.
  • Labeling cost that stays flat as the taxonomy grows. As a product catalog, support taxonomy, or set of monitored risks expands, a supervised system’s data-collection burden grows with every new class. A zero-shot system adds a description.

Challenges of zero-shot classification

The limits are less obvious, and this is where teams get caught.

  • Label descriptions decide accuracy. Vague or overlapping descriptions produce poor results, and there are no labeled examples to fall back on when a description is ambiguous.
  • Pretraining bias carries through. Zero-shot models inherit whatever biases exist in their pretraining data. In hiring, lending, and other sensitive decisions, a biased embedding space can push classifications in a discriminatory direction without an obvious signal. Test outputs across affected groups before deployment and keep testing after it.
  • Scores are hard to explain. There is no fixed decision boundary to point to, only a similarity score between two embeddings or an entailment probability. That makes a specific classification harder to justify to an auditor, a regulator, or a customer who wants a reason. Explainability tooling and clear documentation of labels and thresholds help close the gap.
  • Similarity is not confidence. A cosine distance or an entailment score is not a calibrated probability. Set thresholds on your own validation data, not on intuition.
  • Model changes invalidate stored embeddings. Upgrading the embedding model means re-embedding every label, since vectors from two models do not share a space.

Teams that run zero-shot classification in production treat it as the default for new or rare categories, then add a supervised model or human review for high-stakes decisions where accuracy and explainability both matter.

Zero-shot classification removes the labeled-data bottleneck. It does not remove the need to store, search, and audit the embeddings it depends on. When those embeddings sit in the same SQL database as your application data, every classification is one query away from the records it describes.

To run zero-shot classification over label embeddings in SQL, start a free TiDB Cloud Starter cluster and create the class_labels table above. The free tier includes the vector search features used in this guide.

Zero-Shot Classification FAQ

What is zero-shot classification?

Zero-shot classification is a machine learning technique that assigns input to a category the model was never trained on. Instead of labeled examples, it uses a natural-language description of each category, such as “questions about invoices, refunds, or payment methods.” The model compares the input against each description using knowledge it gained during pretraining and returns the closest match. Adding a new category means writing a new description, not collecting and labeling new training data.

How is zero-shot classification different from supervised learning?

Supervised learning needs labeled examples for every class it recognizes, while zero-shot classification needs only a text description of each class. When enough labeled data exists, a supervised model is more accurate because it learns class boundaries directly from examples. Zero-shot classification wins when categories are new, rare, or change quickly, since a supervised model would need retraining first. Many teams use zero-shot classification for new categories and switch to a supervised model once a category has enough labeled data.

What models are commonly used for zero-shot classification?

Two model families dominate. Natural language inference (NLI) models such as BART-MNLI treat each candidate label as a hypothesis and score how strongly the input supports it, which suits text with a small, fixed set of labels. Embedding models such as CLIP for images and sentence-transformers models like all-MiniLM-L6-v2 for text convert inputs and labels into vectors and compare them. Embedding models scale better to large label sets because label embeddings are computed once and reused.

How do you store zero-shot classification embeddings?

Store zero-shot classification embeddings in a database with vector search. Embed each class-label description once, store the result in a vector column whose dimension matches your model, and add a vector index on cosine distance. To classify new input, embed it and query for the nearest label. In TiDB, that means a VECTOR(384) column for a 384-dimension model and a query that orders by VEC_COSINE_DISTANCE with a LIMIT. This applies to embedding models only, since NLI models produce no reusable label embeddings.

What are the main risks of zero-shot classification?

Zero-shot classification depends on its label descriptions, so vague or overlapping descriptions produce poor results with no labeled data to fall back on. Models also inherit bias from their pretraining data, which matters in sensitive decisions such as hiring or lending. Similarity scores are not calibrated probabilities, so thresholds need tuning on real data, and individual predictions are hard to explain. For high-stakes decisions, pair zero-shot classification with human review or a supervised model.


Last updated 9월 26, 2026

💬 Let’s Build Better Experiences — Together

Join our Discord to ask questions, share wins, and shape what’s next.

Join Now