Enable llama.cpp in Settings
llama.cpp provides efficient CPU and GPU inference for GGUF models. It supports a wide range of quantization formats and is optimized for running large language models on consumer hardware. Features an OpenAI-compatible API.
| Field | Type | Required | Description |
|---|---|---|---|
| baseUrl | url | Optional | URL of the llama.cpp server (llama-server) |
Add provider:
npx webinfer provider add llama-cppTest connection:
npx webinfer provider test llama-cppBasic usage:
import { generateText } from "webinfer"
// WebInfer automatically routes to the best available provider
const result = await generateText({
prompt: "Write a haiku about programming"
})
console.log(result.text)Specify llama.cpp explicitly:
import { generateText } from "webinfer"
const result = await generateText({
prompt: "Write a haiku about programming",
provider: "llama-cpp"
})
console.log(result.text)
console.log("Provider:", result.provider)
console.log("Model:", result.model)Streaming:
import { streamText } from "webinfer"
const { textStream } = await streamText({
prompt: "Write a story about AI",
provider: "llama-cpp"
})
for await (const chunk of textStream) {
process.stdout.write(chunk)
}