Meta logo

Meta: Llama-3.1-8B-Instruct

Compact open Llama model for lightweight chat, drafting, and self-hosting

Source-linked llama Open weights Released 2024-07-23
API record Report
InputT
OutputT
Input price$0.02/M
Output price$0.04/M
Context128K
Max output4.096K
Providers7
Inference availability

Providers

Provider-specific identifiers, limits, and listed prices per million tokens. Every row links back to the provider's own documentation.

Report a price
Providers offering Llama-3.1-8B-Instruct
ProviderProvider model IDContextMax outputInputOutputCache readCapabilitiesDocs
Nvidia meta/llama-3.1-8b-instruct 16K 4.096K Tools Docs ↗
Kilo Gateway meta-llama/llama-3.1-8b-instruct 131.072K 131.072K $0.02 $0.04 ToolsJSON Docs ↗
Inference meta/llama-3.1-8b-instruct 16K 4.096K $0.025 $0.025 Tools Docs ↗
OpenRouter meta-llama/llama-3.1-8b-instruct 131.072K 131.072K $0.05 $0.08 $0.025 ToolsJSON Docs ↗
Hugging Face meta-llama/Llama-3.1-8B-Instruct 131.072K 4.096K $0.06 $0.06 ToolsJSON Docs ↗
Cortecs llama-3.1-8b-instruct 128K 128K $0.167 $0.167 ReasoningToolsJSON Docs ↗
Merge Gateway meta/llama-3.1-8b-instruct 128K 2.048K $0.22 $0.22 Tools Docs ↗

Capability badges appear only where the provider catalog explicitly lists support. A blank cell means the source is silent, not that the feature is absent.

Listed rates

Price across providers

Input price per million tokens as published by each provider. Bars are drawn from listed rates only — no traffic weighting, since the catalog observes no requests.

Lowest input $0.02/M

Across 6 priced providers

Median input $0.055/M

Midpoint of listed rates

Highest input $0.22/M

11.0× the lowest listed rate

Output range $0.025 – $0.22

Per million output tokens

Kilo Gateway $0.02/MLowest
Inference $0.025/M
OpenRouter $0.05/M
Hugging Face $0.06/M
Cortecs $0.167/M
Merge Gateway $0.22/M
Cost calculator

Estimate a workload

$0.00
Excludes taxes, non-token charges, and tiered discounts.

Context limits also differ by provider, from 16K to 131.072K tokens. Compare the provider table above before choosing on price alone.

Specification

Capabilities

Recorded from the source catalog and provider listings.

× Reasoning No
Tool calling Yes
? Structured output Unknown
× Attachments No
× Vision input No
Open weights Yes
Creator
Meta
Model family
llama
Knowledge cutoff
2023-12
License
Not documented
Release date
2024-07-23
Model ID
meta/llama-3.1-8b-instruct

Weights: Hugging Face ↗

Published evaluations

Benchmarks

Every result stays attached to its source, version, metric and harness. Scores from different versions are never merged, and the profile below plots each benchmark against its own population rather than on a shared scale.

Benchmark registry
No published benchmark results

ModelBench does not infer quality from price, context size, or model name. When a source publishes a comparable result, it appears here with its version and link.

Read the methodology
Catalog activity

Change log

Field-level changes detected between successful source imports.

Full change log
No changes recorded

This record has not changed within the retained import history.

Provenance

Sources & verification

Every figure on this page traces back to one of these records.

Methodology
Public API

Use this record

Fetch the complete source-linked model record. No key, no account, no rate-limited tier.

API documentation
Endpoint
GET https://www.modelbench.lol/api/v1/models/meta/llama-3.1-8b-instruct
curl
curl "https://www.modelbench.lol/api/v1/models/meta/llama-3.1-8b-instruct"
Common questions

Frequently asked questions

Answered directly from the stored record — nothing here is generated beyond the catalog's own fields.

What is Llama-3.1-8B-Instruct?

Compact open Llama model for lightweight chat, drafting, and self-hosting. It is published by Meta and catalogued here from Models.dev.

How much does Llama-3.1-8B-Instruct cost?

Listed input pricing starts at $0.02 per million tokens from Kilo Gateway, rising to $0.22 across 6 listed providers.

What is the context length of Llama-3.1-8B-Instruct?

Llama-3.1-8B-Instruct accepts up to 128K tokens of context and returns up to 4.096K output tokens.

Does Llama-3.1-8B-Instruct support tool calling and structured output?

Provider catalogs list support for tool calling.

Which providers serve Llama-3.1-8B-Instruct?

7 providers list this model: Nvidia, Kilo Gateway, Inference, OpenRouter, Hugging Face, Cortecs and 1 more.

Are the weights for Llama-3.1-8B-Instruct open?

Yes. The weights are published and downloadable from Hugging Face.

When was Llama-3.1-8B-Instruct released?

The catalog records a release date of 2024-07-23, last verified Aug 17, 2026.

More models from Meta

View all →