We know small and open-source models. Put that to work.
Beyond the one-time widget, we help teams get real value out of open-weight AI — a quick consultation, high-volume chat for agencies, or a full fixed-cost self-hosted setup for coding assistants, RAG, and automation. Same principle as the widget: no variable token bills.
Three ways to work with us.
Every engagement is priced for what it is — tell us the shape of the problem and we'll come back with a number that makes sense. No calculators, no seat counts.
Consultation
Straight, vendor-neutral guidance on small and open-source models — which one fits your use case, what hardware it needs, and whether to self-host, use a hosted model, or run it browser-native.
- Model selection for your use case — chat, RAG, coding, automation
- Hardware & sizing guidance: what runs on what
- Build-vs-buy and hosted-vs-self-hosted decisions
- A clear recommendation you can act on — no lock-in
Agency & high-volume chat
Deploying Zupport.chat across many brands, or expecting more than 1M widget loads / month on a single assistant? We’ll price it for the volume.
- Higher-volume traffic beyond the 1M / month cap
- Bulk pricing on multiple assistants
- Per-client setup for agencies
- Internal portals across teams or brands
- White-glove onboarding if you need it
- Direct line to the founder, not a ticketing queue
Self-hosted AI, fixed cost
A complete open-source LLM setup for your team — coding assistants, internal chat, RAG, automation — running on a fixed-cost instance you control. No per-token bill, ever.
- We recommend the right open-weight models for the job
- We stand up a private cloud instance / VPS — or hand you an install package
- Hosted and maintained by us, or fully self-owned
- Fixed instance cost — no variable token spend
- Coding assistants, RAG over your docs, workflow automation
- Your data stays on your infrastructure
Typical reply in one business day · straight to a human, no bots.
The same reason our widget has no monthly fee.
Metered AI turns every bit of success into a bigger bill. We don't build that way. The widget runs in the visitor's browser, so there's nothing to meter. A self-hosted setup runs open-weight models on a box you rent, so you pay for the hardware — not per token. Pick a small model that's genuinely good at your task, size the machine right, and the cost stops moving. That's the whole trick, and it's the part most people get wrong. Getting it right is what we're good at.
From "what should we use?" to running it yourself.
Tell us the use case
Coding assistant, support RAG, internal chat, automation — whatever your team actually needs to do.
We recommend models & sizing
The right open-weight model for the workload, and the hardware to run it well. Vendor-neutral.
We deploy it
A private cloud instance or VPS we set up and maintain — or an install package you run yourself.
You run it at a fixed cost
You pay for the box, not per token. Usage grows, the bill stays flat. Your data stays yours.
The honest answers.
Which models do you work with?
Open-weight models like Qwen, Llama, and Mistral, chosen for your use case and hardware. We stay vendor-neutral — the recommendation follows the workload, not a partnership.
Do you host it, or do we?
Either. We can stand up and maintain a private cloud instance or VPS for you, or hand over an install package you run entirely on your own infrastructure.
How is this a fixed cost?
Open-source models run on an instance you rent, so you pay for the box — not per token. Usage goes up, the bill doesn’t. No metered API, no overage bill.
Is my data private?
With a self-hosted setup the models run on your infrastructure. Prompts and documents never leave your environment — nothing is sent to a third-party API.
Do I need this if I just want the chat widget?
No. The one-time $79 Zupport.chat widget is self-contained and needs none of this. Consulting and self-hosted setups are for teams who want coding assistants, RAG, or automation on their own terms.
Tell us what you're building.
A sentence about the use case is enough to start. We'll take it from there.