Common questions.
What is Model Serving?
Model serving is the infrastructure that hosts trained AI models and handles inference requests at scale.
How does Model Serving work?
Serving systems must balance latency, throughput, cost, and availability while supporting governance requirements like logging and access control.
Which terms are related to Model Serving?
Closely related concepts include Inference, Model Endpoint, AI Governance Analyst, AI Proxy. Each is defined in the Prefactor glossary.