Cloudflare Workers AI logo

llama-3.1-8b-instruct-fp8-fast API status on Cloudflare Workers AI

Check provider
★★★★★★★★★★ 3.5 Benchmark-backed score

Llama 3.1 8B Instruct on Cloudflare Workers AI is a lightweight, edge-deployed version of Meta's 8B model optimized with FP8 quantization for fast inference.

Check providerOpenAI compatibleTool callingJSON modeText
API Details

Copy-ready connection details

Use these values in your client or SDK.
Base URL
https://api.cloudflare.com/client/v4/accounts/{account_id}/ai/run
Model ID
cf-meta-llama-3-1-8b-instruct-fp8-fast
API format Cloudflare Workers AI native run + OpenAI Chat Completions
Technical Details

llama-3.1-8b-instruct-fp8-fast specifications

Provider catalog
Context window 131K
Max output 131K
Status Check provider
Family llama-3-1-8b-instruct
Released Jul 23, 2024
Last updated Jun 15, 2026
Free listing since Jul 23, 2024
Input text
Output text
Capabilities tool calling, structured output
AI Recommendation

Should you use llama-3.1-8b-instruct-fp8-fast?

llama-3.1-8b-instruct-fp8-fast is listed for chat workloads and supports a 131K context window.

Check Cloudflare Workers AI's current endpoint status before integrating this listing. Use the alternatives table if the provider listing is unavailable.

Best for
  • Chat
Strengths & Weaknesses

Strengths

  • Long context window
  • Tool calling support
  • Structured JSON output

Watch outs

  • Current endpoint availability needs provider confirmation
  • Free-tier rate limits apply
  • Vision support is not listed
Benchmark Overview

Benchmark signals for llama-3.1-8b-instruct-fp8-fast

Measured data
Intelligence General reasoning and instruction following
7.6/100
Coding Programming and code generation
5.4/100
Agentic Tool use and multi-step tasks
0.5/100
Speed Observed generation speed
166 tok/s
Context Maximum listed context window
131K
Pricing

llama-3.1-8b-instruct-fp8-fast pricing per 1M tokens

Check provider
Input $0.1 per 1M tokens
Output $0.1 per 1M tokens
Free access Check provider Cloudflare Workers AI
Rate limit 10K neurons/day (shared) provider policy
Availability

llama-3.1-8b-instruct-fp8-fast availability by provider

2 alternatives

We found 3 provider listings for llama-3-1-8b-instruct. Check model ID, quota, pricing, and API format before switching providers.

Provider Model listing Access Context API Limits
Cloudflare Workers AI @cf/meta/llama-3.1-8b-instruct-fp8-fast Check provider 131K Native 10K neurons/day (shared)
Hugging Face Meta-Llama-3.1-8B-Instruct Free tier 128K Native Credit-metered
NVIDIA NIM meta/llama-3.1-8b-instruct Free tier 8K Native Varies
View Cloudflare Workers AI setup guide →
Typical Use Cases

llama-3.1-8b-instruct-fp8-fast use cases

Chat

llama-3.1-8b-instruct-fp8-fast is tagged for chat in this catalog and works with OpenAI-compatible client libraries.

FAQ

llama-3.1-8b-instruct-fp8-fast free API FAQ

Is llama-3.1-8b-instruct-fp8-fast free to use?

llama-3.1-8b-instruct-fp8-fast appears in the free model catalog for Cloudflare Workers AI, but its current endpoint availability should be confirmed with the provider before use.

What is the llama-3.1-8b-instruct-fp8-fast model ID?

The model ID shown in this catalog is cf-meta-llama-3-1-8b-instruct-fp8-fast.

What are the llama-3.1-8b-instruct-fp8-fast free tier rate limits on Cloudflare Workers AI?

The listed free tier limit is 10K neurons/day (shared). Limits can change per account tier, so confirm against the provider dashboard.

What context window does llama-3.1-8b-instruct-fp8-fast support?

The listed context window is 131K tokens with up to 131K output tokens.

More about llama-3.1-8b-instruct-fp8-fast

Llama 3.1 8B Instruct on Cloudflare Workers AI is a lightweight, edge-deployed version of Meta's 8B model optimized with FP8 quantization for fast inference. It is the fastest and most cost-efficient option in Cloudflare's free catalog — ideal for high-volume, low-latency tasks like chat routing, text classification, summarization, or simple Q&A where a 70B+ model would be overkill. Like all Cloudflare Workers AI models, it shares the 10,000 Neurons/day free pool and uses a Cloudflare-specific API format. For developers already on Cloudflare's ecosystem, it integrates directly with Workers and Pages with minimal cold start.

For API keys, setup steps, and provider-level limits, see the Cloudflare Workers AI provider page.