Cloudflare Workers AI logo

llama-3.3-70b-instruct-fp8-fast Free API on Cloudflare Workers AI

Free API Verified
★★★★★★★★★★ 3.5 Benchmark-backed score

Popular open Llama workhorse for multilingual chat, coding, and self-hosting

Free APIOpenAI compatibleTool callingJSON modeText
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run
Model ID
@cf/meta/llama-3.3-70b-instruct-fp8-fast
API format Cloudflare Workers AI native run + OpenAI Chat Completions
Technical Details

llama-3.3-70b-instruct-fp8-fast specifications

Provider and model catalog
Context window 131K
Max output 131K
Status Online
Family llama
Knowledge cutoff 2023-12
Released Dec 6, 2024
Last updated Aug 6, 2026
Free listing since Dec 6, 2024
Input text
Output text
Capabilities tool calling, structured output, file attachments, temperature control
Open weights Yes

External benchmark references

Benchmark Score Metric Date
Artificial Analysis Coding Index 10.7 index 2026-03-11
SciCode 26 percent correct 2026-03-11
Terminal-Bench Hard 3 success rate 2026-03-11
AI Recommendation

Should you use llama-3.3-70b-instruct-fp8-fast?

llama-3.3-70b-instruct-fp8-fast is listed for chat workloads and supports a 131K context window.

Use it when Cloudflare Workers AI's free tier is enough for evaluation, demos, or light production traffic.

Best for
  • Chat
Strengths & Weaknesses

Strengths

  • Long context window
  • Tool calling support
  • Structured JSON output
  • Open weights available

Watch outs

  • Free-tier rate limits apply
  • Vision support is not listed
Benchmark Overview

Benchmark signals for llama-3.3-70b-instruct-fp8-fast

Measured data
Intelligence General reasoning and instruction following
9.4/100
Coding Programming and code generation
11.9/100
Agentic Tool use and multi-step tasks
0.3/100
Speed Observed generation speed
80 tok/s
Context Maximum listed context window
131K
Pricing

llama-3.3-70b-instruct-fp8-fast pricing per 1M tokens

Free tier listed
Input $0.6 per 1M tokens
Output $0.72 per 1M tokens
Free access Available Cloudflare Workers AI
Rate limit 10K neurons/day (shared) provider policy
Availability

llama-3.3-70b-instruct-fp8-fast availability by provider

3 alternatives

We found 4 provider listings for llama-3-3-70b-instruct. Check model ID, quota, pricing, and API format before switching providers.

Provider Model listing Access Context API Limits
Cloudflare Workers AI @cf/meta/llama-3.3-70b-instruct-fp8-fast Free tier 131K Native 10K neurons/day (shared)
GitHub Models Llama-3.3-70B-Instruct Free tier 131K OpenAI-style 15 RPM, 150 RPD
OVHcloud AI Endpoints Meta-Llama-3_3-70B-Instruct Free tier 131K OpenAI-style 2 RPM (anonymous)
NVIDIA NIM Llama-3.3-70B-Instruct Free tier 128K Native Varies
View Cloudflare Workers AI setup guide →
Typical Use Cases

llama-3.3-70b-instruct-fp8-fast use cases

Chat

llama-3.3-70b-instruct-fp8-fast is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

FAQ

llama-3.3-70b-instruct-fp8-fast free API FAQ

Is llama-3.3-70b-instruct-fp8-fast free to use?

llama-3.3-70b-instruct-fp8-fast is listed with free API access on Cloudflare Workers AI, subject to the provider's quota and account policy.

What is the llama-3.3-70b-instruct-fp8-fast model ID?

The model ID shown in this catalog is @cf/meta/llama-3.3-70b-instruct-fp8-fast.

What are the llama-3.3-70b-instruct-fp8-fast free tier rate limits on Cloudflare Workers AI?

The listed free tier limit is 10K neurons/day (shared). Limits can change per account tier, so confirm against the provider dashboard.

What context window does llama-3.3-70b-instruct-fp8-fast support?

The listed context window is 131K tokens with up to 131K output tokens.

More about llama-3.3-70b-instruct-fp8-fast

Llama 3.3 70B Instruct runs on Cloudflare Workers AI's global edge network, bringing Meta's 70B-parameter flagship to every Cloudflare data center worldwide. Deployed at the edge, it offers significantly lower latency than centralized API providers — requests are routed to the nearest Cloudflare PoP rather than a single-region GPU cluster. The free tier allocates 10,000 Neurons (compute units) per day across all models on your account, so available capacity depends on your other Workers AI usage. The API uses a Cloudflare-specific REST format rather than OpenAI SDK compatibility, and requires a Cloudflare account ID in the endpoint URL — lightweight setup for existing Cloudflare users, but an extra step for everyone else.

For API keys, setup steps, and provider-level limits, see the Cloudflare Workers AI provider page.