Bring Your Own Model

View as Markdown

For privacy, cost control, or specialized models, you can run LLMs locally or on your own servers. The OpenAIClient works with any endpoint that implements the OpenAI chat completions API.

A localhost base URL (like the Ollama, vLLM, and LM Studio examples below) only works when you run the crew locally (python server.py + smallestai agent-crew chat). A deployed crew runs in the cloud and cannot reach localhost on your machine. For a deployed agent, point base_url at a host the cloud can reach: a public inference endpoint, or your local server exposed through a tunnel. For Ollama over a tunnel, rewrite the Host header or Ollama returns 403: ngrok http 11434 --host-header=localhost:11434, then set base_url="https://<id>.ngrok-free.app/v1".

Tunnels are for development only. If the local process behind the tunnel goes down (Ollama exit, laptop sleep), the tunnel returns ERR_NGROK_8012 and the agent surfaces it as a fatal agent_error and hangs up mid-call. Move to a hosted OpenAI-compatible endpoint before shipping.

Do not point base_url at Anthropic’s OpenAI-compatible endpoint (https://api.anthropic.com/v1/) for a crew agent. It truncates streaming replies and drops tool-call turns, which surfaces as the agent going silent mid-conversation. If you want Claude for a crew, route it through a gateway that stabilizes the stream (LiteLLM, OpenRouter, Portkey), use a real OpenAI-backed model, or use the platform model. Anthropic’s native SDK is unaffected. Only their OpenAI-compat URL has this issue.

Complete Example

Here’s a full agent using a local Ollama model:

1import os
2from smallestai.atoms.crew.nodes import OutputCrewNode
3from smallestai.atoms.crew.clients.openai import OpenAIClient
4from smallestai.atoms.crew.server import AtomsCrewApp
5from smallestai.atoms.crew.session import CrewSession
6
7class LocalAgent(OutputCrewNode):
8 def __init__(self):
9 super().__init__(name="local-agent")
10
11 # Connect to your local model
12 self.llm = OpenAIClient(
13 model="llama3",
14 base_url="http://localhost:11434/v1",
15 api_key="ollama" # Not required for Ollama
16 )
17
18 self.context.add_message({
19 "role": "system",
20 "content": "You are a helpful assistant running on local hardware.",
21 })
22
23 async def generate_response(self):
24 response = await self.llm.chat(
25 messages=self.context.messages,
26 stream=True
27 )
28 async for chunk in response:
29 if chunk.content:
30 yield chunk.content
31
32async def on_start(session: CrewSession):
33 session.add_node(LocalAgent())
34 await session.start()
35 await session.wait_until_complete()
36
37if __name__ == "__main__":
38 app = AtomsCrewApp(setup_handler=on_start)
39 app.run()

Requirements

Your model server must implement the OpenAI Chat Completions API:

FeatureRequiredNotes
/chat/completions endpointYesStandard OpenAI format
StreamingYesstream=True must work
Tool callingFor toolsOpenAI-format function calling

Custom Endpoints

Connect to any custom model server:

1from smallestai.atoms.crew.clients.openai import OpenAIClient
2
3llm = OpenAIClient(
4 model="your-model-name",
5 base_url="https://your-server.example.com/v1",
6 api_key=os.getenv("YOUR_API_KEY")
7)

Ollama

Ollama is the easiest way to run models locally. It handles model downloads and serving automatically.

Setup

$# Install Ollama
$curl -fsSL https://ollama.com/install.sh | sh
$
$# Pull a model
$ollama pull llama3
$
$# Start the server (runs on port 11434)
$ollama serve

Usage

1from smallestai.atoms.crew.clients.openai import OpenAIClient
2
3llm = OpenAIClient(
4 model="llama3",
5 base_url="http://localhost:11434/v1",
6 api_key="ollama" # Doesn't require a real key
7)

vLLM

vLLM is a high-performance inference server for production workloads.

Setup

$pip install vllm
$
$python -m vllm.entrypoints.openai.api_server \
> --model meta-llama/Llama-3-8B-Instruct \
> --port 8000

Usage

1from smallestai.atoms.crew.clients.openai import OpenAIClient
2
3llm = OpenAIClient(
4 model="meta-llama/Llama-3-8B-Instruct",
5 base_url="http://localhost:8000/v1",
6 api_key="vllm"
7)

LM Studio

LM Studio provides a desktop UI for running models locally.

  1. Download from lmstudio.ai
  2. Load a model
  3. Start the local server (Settings → Local Server)
1from smallestai.atoms.crew.clients.openai import OpenAIClient
2
3llm = OpenAIClient(
4 model="local-model",
5 base_url="http://localhost:1234/v1",
6 api_key="lmstudio"
7)

Deterministic tool calls with tool_choice

Crew OpenAIClient.chat() forwards extra kwargs to the underlying request, so any OpenAI-compatible tool_choice value works. Passing tool_choice="required" on a specific turn forces the model to emit a tool call rather than a free-text answer. Useful when the model has already said “let me transfer you” out loud but does not actually invoke the transfer tool.

1response = await self.llm.chat(
2 messages=self.context.messages,
3 tools=self.tools,
4 stream=True,
5 tool_choice="required", # this turn must produce a tool call
6)

Set tool_choice="required" only for the specific transfer or handoff turn, not on every turn. Forcing tools on normal conversational turns will cause the model to invent tool calls when the user is just chatting.

Troubleshooting

IssueCauseFix
Connection refusedServer not runningStart Ollama/vLLM
Model not foundWrong nameCheck ollama list or server logs
No streamingServer configEnsure streaming is enabled
Tool calls ignoredModel limitationUse a larger model, add tool_choice="required" for the decisive turn, or use a cloud fallback
Agent hangs up mid-call with agent_errorTunnel returned ERR_NGROK_8012 because the local model process diedRestart the local process, or move to a hosted OpenAI-compatible endpoint before shipping

Tips

Ollama is great for local development. For production, consider vLLM or a cloud provider for reliability.

Local models can fail. Configure a cloud fallback (e.g., OpenAI) to catch errors and keep the conversation going.

Not all local models support function calling. Test your tools or use a model known to support them.