Pricing
Free
Get started with Ollama
- Automate coding, document analysis, and other tasks with open models
- Keep your data private
- Run models on your hardware
- Access cloud models
- CLI, API, and desktop apps
- 40,000+ community integrations
- Unlimited public models
Pro
Solve harder tasks, faster
or $200/yr billed annually
- Access larger, more powerful cloud models
- Run 3 cloud models at a time
- 50x more cloud usage than Free
- Upload and share private models
Max
For your most demanding work
New Max subscriptions are temporarily paused while we add capacity. Learn more
- Run 10 cloud models at a time
- 5x more usage than Pro
Team
Introductory pricingBring open models to your whole team.
5-seat minimum, usage included
- Access to powerful open models in the US and Europe
- High performance, up to 2x more than model gateways
- Zero data retention and logging
- Shared billing and administration
- Priority support
- Single sign-on (SSO)
- Model access controls
- MDM installer for Windows and macOS
Enterprise
Custom terms and support for larger organizations.
Talk with our team about your needs
- Volume pricing and custom terms
- Security and procurement support
- Deployment planning with Ollama
Frequently asked questions
Plans
-
Why are new Max subscriptions paused?
Ollama's cloud has more than doubled in token volume every month, and with even larger open models like kimi-k3 coming soon, demand is growing faster than we can add capacity. Max subscribers run some of the heaviest workloads, so new Max subscriptions are paused while we add more capacity to protect the experience of existing subscribers.
Existing Max subscribers keep their plan, limits, and pricing. Pro and Free remain open, and Pro subscribers can add extra usage to go beyond their plan's included limits.
Team
-
Who is the Team plan for?
The Team plan is for groups that want shared billing for seats and model usage.
-
How does Team billing work?
Teams start at five seats. Each seat costs $25 per month, so the minimum seat charge is $125 per month. Each additional seat adds $25 per month.
Each seat includes usage. Usage beyond that draws from a shared team balance billed as you go. You can add a set amount to the balance and turn off automatic usage billing.
-
Does joining a team change my personal account?
Ollama creates a separate team account for you when you join. You can switch between your team and personal accounts with the same login.
Models
-
Which models are available?
See the full list of cloud-enabled models here.
-
Do models support tool calling?
Yes. Cloud models that are trained to support tools are tested for tool calling and with real agent workflows before they go live. If something isn't working, let us know at [email protected].
-
What quantization or data format do cloud models use?
Native weights, as released by the model provider. On modern NVIDIA hardware, models may use accelerated data formats supported by Blackwell and Vera Rubin architectures (e.g. NVFP4).
-
How fast is Ollama?
Speed depends on model size, architecture, and hardware optimization. We target and monitor for low time-to-first-token and high throughput across all cloud models. Priority tiers with faster performance may be available in the future.
Usage
-
What are the usage limits for each plan?
Running models on your own hardware is always unlimited. Cloud usage varies by plan:
Plan Usage Example use cases Free Light usage Chatting with models, evaluating larger models, coding and AI assistants with smaller models Pro Day-to-day work Larger models, coding automation, deep research Max Heavy, sustained usage Continuous agent tasks, multiple concurrent agents, large models over extended sessions Each plan has session limits that reset every 5 hours and weekly limits that reset every 7 days.
-
How is usage measured?
Individual plans have usage limits based on the model and the number of input, cached input, and output tokens processed. They don't cap you at a fixed number of tokens because different models use different amounts of compute.
For teams, each member's usage draws from the usage included with their seat first. Once it's used, further usage draws from the team's shared extra usage balance at the model's token rate.
-
How much usage does each model use?
Models consume a different amount of usage based on how difficult they are to run. To view a model's usage level, visit the model's page, where its usage level is displayed from small, light models (level 1), like gpt-oss:20b, to extra heavy models (level 4), like deepseek-v4-pro.
-
How does extra usage work?
Pro and Max users can add extra usage balance. Ollama uses included plan limits first, then draws from the extra usage balance. Team usage draws from one balance shared by the organization.
-
How much more usage does Pro include?
50x more than Free.
-
How much more usage does Max include?
5x more than Pro.
-
How do I know when I've hit my limit?
Check your usage here anytime. At 90% of your plan's limit, Ollama sends an email reminder. You can turn this off in settings.
-
How many cloud models can I run at once?
Concurrency limits ensure dedicated capacity for workflows that need multiple models running simultaneously:
Plan Concurrent models Free 1 Pro 3 Max 10 Requests beyond your plan's concurrency limit are queued and processed as soon as a slot is available. Queued requests are held up to a fixed limit - if the queue is full, the request will be rejected until one of your concurrency slots opens.
Accounts
-
Can I have multiple Ollama accounts?
No. Ollama is one account per person.
Privacy
-
Where are models hosted?
Ollama hosts models and compute resources primarily in the United States. To serve global demand, we may route to Europe and Singapore for additional capacity.
-
Is my prompt or response data trained on?
Prompt or response data is never logged or trained on.
-
Who does Ollama partner with to host models?
Ollama collaborates with NVIDIA Cloud Providers (NCPs) to host open models.
When Ollama partners with providers, we require no logging, no training, and zero data retention policies in place.