Open Weight · Format Explorer

Safetensors, GGUF and the model format layer.

Understand what a weight format is, how tensors and metadata are stored, why checkpoints are sharded, where quantization fits, and how formats connect to runtimes and hardware.

SafetensorsSafe, efficient tensor serialization
GGUFModel metadata + tensors for GGML workflows
ShardingLarge checkpoints split across files
CompatibilityFormat + runtime + architecture
01 · Foundation

What is a model weight format?

A model weight format defines how learned tensors — and sometimes associated metadata — are serialized into files that software can store, distribute and load.

Important: a weight format is not the same thing as a model architecture, numerical precision, quantization method or inference runtime.
Architecture
≠
Weight format
≠
Precision
≠
Quantization
≠
Runtime
≠
Hardware
02 · Safetensors vs. GGUF

Two formats, different design contexts

DimensionSafetensorsGGUF
Core roleTensor serialization formatGGML model file format containing metadata and tensors
Typical ecosystemHugging Face / Transformers and many ML workflowsllama.cpp / GGML-oriented inference workflows
Safety designDesigned to avoid arbitrary code execution associated with pickle-style loadingBinary structured model format; security still depends on the complete loading stack
MetadataSupports metadata alongside tensor informationUses structured key-value metadata plus tensor descriptors
Quantized modelsCan be used in multiple quantized workflows depending on framework/runtimeExplicitly supports many quantized GGML tensor types
Local inferencePossible with compatible runtimesCommonly used for local llama.cpp inference
Server inferenceCommon across server-oriented ML stacksPossible in supported runtimes; support and optimization vary

This comparison describes common ecosystem usage, not a universal rule. Runtime support changes over time.

03 · Safetensors

Safe and efficient tensor storage

Hugging Face describes Safetensors as a simple format for storing tensors safely while supporting fast loading, including zero-copy behavior in supported workflows.

It is widely used for model checkpoints distributed through the Hugging Face ecosystem.

Tensor serialization Metadata Fast loading HF ecosystem
model.safetensors ┌────────────────────────────┐ │ Header / tensor metadata │ ├────────────────────────────┤ │ Tensor data │ │ ... │ └────────────────────────────┘

Official Safetensors documentation ↗

04 · GGUF

Metadata + tensors for GGML workflows

Hugging Face documents GGUF as a single-file format used to store models for inference with GGML, containing both model metadata and tensors. It supports many quantized data types.

The canonical llama.cpp implementation defines a GGUF structure with a header, key-value metadata, tensor descriptors and tensor data.

GGML llama.cpp Metadata Quantized types
GGUF ┌────────────────────────────┐ │ Magic + version │ ├────────────────────────────┤ │ Key-value metadata │ ├────────────────────────────┤ │ Tensor descriptors │ ├────────────────────────────┤ │ Tensor data │ └────────────────────────────┘

Transformers GGUF documentation ↗

llama.cpp ↗

05 · Sharding

Why Safetensors checkpoints may use many files

Very large checkpoints are frequently sharded: the tensors are distributed across multiple files, with an index describing which tensor belongs to which shard.

model-00001-of-00004.safetensors model-00002-of-00004.safetensors model-00003-of-00004.safetensors model-00004-of-00004.safetensors model.safetensors.index.json
Sharding does not change the model architecture. It changes how a large checkpoint is packaged and loaded.
06 · Metadata

Weights need context

A collection of tensors is not enough for a reliable deployment. Software also needs enough information to interpret those tensors correctly.

architecture
How tensor names and shapes map to model structure.
tokenizer
How text or other inputs are converted into model tokens.
dtype
The numerical representation of stored parameters.
quantization
How lower-precision representations were produced.
revision
Which source checkpoint or repository state produced the artifact.
license
What use and redistribution terms apply.

Some metadata lives inside the weight file; some lives in adjacent configuration, tokenizer or repository files. The exact arrangement depends on format and ecosystem.

07 · Quantization

Format and quantization are different layers

Quantization changes how values are represented. A file format determines how data is serialized. They are related but not interchangeable concepts.

Model tensors
→
Precision / quantization
→
Serialization format
→
Runtime loader
→
Hardware

For example, GGUF supports many quantized GGML data types, while Safetensors can store tensors used by multiple framework-specific quantization approaches.

08 · Conversion

What happens when a model is converted?

Source checkpoint ↓ Read configuration + tokenizer + tensors ↓ Map tensor names / structure ↓ Optional quantization ↓ Write target format + metadata ↓ Validate output ↓ Run inference checks

llama.cpp provides conversion tooling for supported architectures, including workflows that convert Hugging Face checkpoints into GGUF.

Conversion is not just renaming a file extension. Architecture support, tensor mapping, tokenizer metadata and quantization choices must all be compatible.
09 · Compatibility

The format is only useful if the runtime can load it

Runtime / workflowSafetensorsGGUFImportant nuance
TransformersNative ecosystem fitSupported for documented GGUF loading workflowsFeature support depends on architecture and current Transformers implementation.
llama.cppConversion typically requiredPrimary formatllama.cpp requires supported models to be in GGUF for its standard model-loading workflow.
vLLMCommon loading pathExperimental / plugin-basedCurrent vLLM documentation describes GGUF support as highly experimental and under-optimized.
Local desktop workflowsDepends on application/runtimeCommon in llama.cpp-based toolsApplication support matters more than the filename alone.

Compatibility should always be checked against the current runtime version and the specific model architecture.

10 · Decision helper

Which direction fits your workflow?

Choose a scenario. The answer is a workflow starting point — not a universal rule.

Start with Safetensors.

It is a natural fit for many Hugging Face / Transformers checkpoint workflows. Then choose precision or quantization according to your framework and deployment target.

11 · Common mistakes

Five concepts people often mix up

Format ≠ architecture

A Transformer architecture can be distributed in different weight formats.

Format ≠ quantization

INT4 describes numerical representation, not by itself a universal file format.

Weights ≠ tokenizer

Tokenization artifacts are separate from the learned tensor values.

Extension ≠ compatibility

A runtime also needs support for the architecture and tensor representation.

Conversion ≠ zero risk

Converted artifacts should be validated against the source model.

Open weights ≠ open source

Weight access does not automatically open training data, code or usage rights.

12 · Quick reference

Format-selection checklist

QuestionWhy it matters
Which runtime will load the model?Runtime support determines which formats and architectures are practical.
Which hardware is the target?CPU, GPU and Apple Silicon workflows can favor different runtimes and quantizations.
Is this for training or inference?Training, fine-tuning and optimized local inference have different needs.
Do you need quantization?Quantization affects memory, quality and compatibility.
Is the model architecture supported?A valid file format does not guarantee runtime support for the architecture.
Can the artifact be traced to its source?Provenance matters for reproducibility and security.
Next

Continue the Open Weight series

Open Weight Explorer

The fundamentals of tensors, weights, precision and deployment.

Quantization Explorer

FP16, BF16, FP8, INT8, INT4 and practical trade-offs.

Model Portability Explorer

Formats, runtimes, hardware and compatibility.

Primary sources

Technical references

Hugging Face — Safetensors ↗

Hugging Face Transformers — GGUF ↗

ggml-org / llama.cpp ↗

vLLM — GGUF ↗