← Back to glossaryGlossary

Model Serving

Reviewed 19 July 2026Canonical definitionPart of: Agent Observability Terms →

Model serving is the infrastructure that hosts trained AI models and handles inference requests at scale. Serving systems must balance latency, throughput, cost, and availability while supporting governance requirements like logging and access control.

§01 / QUESTIONSterm: Model Serving
Questions

Common questions.

What is Model Serving?

Model serving is the infrastructure that hosts trained AI models and handles inference requests at scale.

How does Model Serving work?

Serving systems must balance latency, throughput, cost, and availability while supporting governance requirements like logging and access control.

Which terms are related to Model Serving?

Closely related concepts include Inference, Model Endpoint, AI Governance Analyst, AI Proxy. Each is defined in the Prefactor glossary.

§02 / RELATEDnext: where this fits
Keep reading

Where this fits.

See how every agent performs, and make it better

Prefactor helps teams observe, evaluate, and improve their AI agents in production, across every framework and provider.