Request Callback
AI Inference as a Service

AI inference APIs, with zero infrastructure to manage

GPU Virtual Machine NVIDIA

Overview of AI Inference

AI Inference as a Service is a managed, CloudXP marketplace-based platform that gives your teams on-demand access to multiple enterprise-grade open-weight large language models through a single, standard REST API. You subscribe to the model that fits your use case, and the platform handles provisioning, hosting, scaling, and security.

No GPU procurement, no model deployment pipeline, and no infrastructure team required. The service exposes a standard chat completions endpoint. Your application sends structured prompts in and receives structured or natural-language outputs out. Everything from prompt testing to API key management is handled through the AI Inference console, which is part of the Jio CloudXP portal. A built-in Playground lets teams design and validate prompts before writing a single line of integration code.

Variants of large language models

Multiple enterprise LLMs, one catalogue
Choose the model that matches your task. Each model is available on a Free Tier for development and evaluation.

Llama 3.1 70B

Best for - General-purpose chat and reasoning

Parameters - 70 Billion

Environment - RAG, Production chat, summarisation

Llama 3.1 405B

Best for - Advanced reasoning and complex instructions

Parameters - 405 billion

Environment - Frontier reasoning, complex multi-step tasks

Qwen 3 30B

Best for - Multilingual tasks and instruction following

Parameters - 405 billion

Environment - Frontier reasoning, complex multi-step tasks

Qwen 3 235B

Best for - Large-context, high-accuracy generation

Parameters - 235 billion

Environment - Long-context generation, high-accuracy output

DeepSeek V3 671B

Best for - Code generation and deep reasoning

Parameters - 671 billion

Environment - Code generation, deep technical reasoning

DeepSeek R1 70B

Best for - Research-grade reasoning and analysis

Parameters - 70 billion

Environment - Research analysis, structured reasoning

Kimi 2.7

Best for - Long-horizon coding, agentic tool workflows, multi-step software engineering tasks, and native multimodal (vision) tasks

Parameters - 1 Trillion

Environment - Long-horizon agent runs, multi-file codebases, complex debugging

Core features of AI Inference

Curated model catalogue
Multiple enterprise-grade open-weight LLMs in a single catalogue: Kimi, Llama, Qwen, and DeepSeek families, each purpose-built for a different class of task. Browse, compare, and order directly from the Marketplace.
Instant provisioning
Place a Marketplace order and the platform provisions a dedicated instance automatically. Five milestone checkpoints give real-time visibility into the provisioning pipeline. Typically active in minutes.
Built-in Playground
Test and iterate on prompt designs directly in the browser: no code, no SDK. Set system prompts, temperature, and token limits. Monitor daily token usage in real time before writing any integration code.
Standard chat completions API
A single POST endpoint per instance following the standard chat completions format. Structured prompts in, structured or natural-language outputs out. Ready-made curl commands provided in the portal.
Predictable tier-based pricing
Fixed tier pricing with defined rate limits and daily token allocations: no surprise usage bills. A Free Tier is available for all models for development and evaluation. Upgrade by provisioning a new instance at a higher tier.
Enterprise security
TLS 1.2+ in transit, AES-256 at rest. Request data is not persisted after each call. API keys are managed exclusively through the console and never embedded in responses.

What you get with AI Inference

Frequently asked questions on AI Inference

Multiple enterprise-grade open-weight models are available through the Marketplace: Llama 3.1 70B (general-purpose chat and reasoning), Llama 3.1 405B (advanced reasoning and complex instructions), Qwen 2.5 72B (multilingual tasks and instruction following), Qwen 3 235B (large-context, high-accuracy generation), DeepSeek V3 671B (code generation and deep reasoning), Kimi 2.7 and DeepSeek R1 70B (research-grade reasoning and analysis). All multiple models follow the same API format and are accessible through the same Marketplace and console workflow.
The complete workflow is four steps: (1) In the Jio CloudXP portal, go to Marketplace, then Products and Templates, then Product, and search for 'llm'. (2) Select a model, complete the order form (project, subscription, location, tier), and confirm. (3) Wait for all five provisioning milestones to complete in Order History. (4) Navigate to Resources, then AI Services, then AI Inference (Text to Text), click your instance's Asset ID, and open the Playground tab or API Configuration tab to start calling the API. The Free Tier is available for all models and requires no billing setup for development and evaluation use.
API keys are managed from the AI Inference console. Select your instance row in the instance table and use the toolbar buttons. Three operations are available: Create (generates a new API key, shown once in the dialog — copy and store immediately in a secrets manager or environment variable, as the portal will never show the full key again), Rotate (issues a new key and immediately invalidates the current key in a single atomic step — update all applications before rotating), and Revoke (disables the current key immediately — all API calls with the revoked key return 401 Unauthorised). One active key per instance at all times. Never hardcode API keys in source code or include them in client-side bundles.
No. The tier is fixed at provisioning time and cannot be changed on an existing instance. To access a higher tier, provision a new instance at the required tier through the Marketplace. You can have multiple instances of the same model simultaneously, for example one Free Tier instance for development and one paid-tier instance for production. Rate limits and token ceilings for each tier are visible in the Base Configuration tab of each instance.
The service applies enterprise-grade security controls end-to-end: all data in transit is encrypted with TLS 1.2 or higher, data at rest uses AES-256 encryption, request payloads are not stored or persisted after processing, no training is performed on customer data, and API keys are created, rotated, and revoked from the console with full lifecycle control.

Ready to build with enterprise LLMs

Provide your details and our team will reach out with model selection guidance, pricing, and a hands-on demo of the Playground and API integration.
image
Captcha
By selecting 'Submit', you authorise Jio Platforms Limited to store your contact details for further communication.
Submit
Cancel