A Standard for Discoverable Local AI Services
One service on your local computer or on a home or office network can serve several AI applications. Serving LLMs, text and image embedding models, speech transcription, image generation and much more.
But local AI apps still need too much setup: a server address, port, model name, and knowledge of what each model supports.
A standard should answer two questions: Where is the service? What can its models do?
Keep an OpenAI-compatible API for inference, then add service discovery and per-model architecture metadata. The names and metadata below are proposals, not an existing standard.
Discover model inputs and outputs
The OpenAI models endpoint lists model IDs and basic metadata. Extend /v1/models with an architecture object for each model.
OpenRouter's model catalog already uses the architecture object with input_modalities and output_modalities. Build on that idea with embedding-specific metadata:
{
"id": "siglip-clip-text-vision",
"architecture": {
"input_modalities": ["text", "image"],
"output_modalities": ["embedding"],
"embedding": {
"dimensions": 768,
"shared_space": true
}
}
}
An embedding turns input into a vector for tasks such as similarity search. This example accepts text and images and returns vectors with 768 dimensions.
shared_space: true means this model maps its input modalities into the same vector space. A photo app could compare a text query with image embeddings directly.
A text search app could select models with text input and embedding output. A photo app could require image input too. Applications could choose suitable models without hard-coding their names.
The nested embedding object is a proposed extension to the OpenRouter-inspired structure. Other model details, such as tool support, can live in separate metadata fields.
Keep familiar inference routes:
POST /v1/chat/completions
POST /v1/embeddings
POST /v1/audio/transcriptions
POST /v1/audio/speech
The standard must also define how these inputs and outputs map to request and response formats. Image and audio embeddings need agreed extensions beyond the text input format.
Find the server automatically
A fixed port does not solve discovery. Ports can conflict, and users still need the server's address.
Use DNS Service Discovery (DNS-SD) instead. On a local network, it can run over Multicast DNS (mDNS), as described in RFC 6763 and RFC 6762.
A server could advertise this proposed service type:
Service: _local-ai._tcp.local.
Instance: Desktop AI
Host: workstation.local.
Port: 43821
DNS-SD lets clients browse service instances and resolve each hostname and port. The server can choose any available port.
Keep discovery records small. Put detailed metadata in HTTP. After discovery, the client could request GET /.well-known/ai:
{
"protocol": "local-ai",
"version": "1",
"api": {
"type": "openai-compatible",
"base": "/v1"
}
}
The standard would also need to specify transport and authentication, so clients know how to connect before fetching this document.
The full flow stays small:
DNS-SD / mDNS: _local-ai._tcp.local.
↓ hostname + port
GET /.well-known/ai
↓ protocol + API base
GET /v1/models
↓ models + architecture metadata
POST /v1/*
↓ inference
Managed networks can publish DNS-SD records through unicast DNS in a configured browsing domain. This supports discovery across routed networks where local mDNS does not reach.
DNS-SD finds the service. The discovery document identifies the protocol. Model architecture describes inputs and outputs. The inference API runs the request.
Together, these pieces could let applications share local AI infrastructure without asking users to manage server addresses, ports, or model-specific setup.