Skip to main content
Cloud & AI Hub
Browse
Glossary AI Directory Playgrounds Models Prompts Explainers Strategy Matrix Benchmark Decoder
Microsoft Released: 2024-05-21

Phi-3 Small

Model Specifications

Context Window 128k tokens
Parameters 7B
Pricing (Input) $0.20 / M tokens
Pricing (Output) $0.60 / M tokens

What is Phi-3 Small?

Phi-3 Small is a 7 billion parameter model in Microsoft’s Phi-3 family, released in May 2024. It utilizes a block-diagonal attention mechanism to speed up token generation.

Like its sibling models, it is trained on curated synthetic data and high-quality web sources, providing a balanced profile for mid-tier tasks.

Key Capabilities

  • Optimized attention: Block-diagonal attention reduces processing times.
  • MIT License: Open for commercial deployment.
  • Balanced profile: Strong reasoning relative to its footprint.

Ideal Use Cases

  • Local assistants: Powering coding and drafting assistants on work laptops.
  • Data extraction: Parsing specific fields from unstructured text.
  • API orchestration: Translating inputs to API routing commands.

Limitations & Caveats

  • Synthetic-data training trade-off: Phi-3 models are trained heavily on curated, textbook-quality and synthetic data to maximize benchmark performance per parameter — this can produce narrower general-world knowledge than models trained on more diverse web-scale corpora of similar size.
  • Superseded within Microsoft’s own lineup: Later Phi-3.5 and Phi-4 models improve on Phi-3 Small’s reasoning and instruction-following at similar or smaller sizes.
  • Weaker code generation than size-matched rivals: A HumanEval score of 52.4 trails contemporaries like Qwen 2.5 7B, making it a better fit for lightweight reasoning and drafting tasks than as a primary coding assistant.

Phi-3 Small’s Block-Diagonal Attention Trade-off

The block-diagonal attention mechanism used in Phi-3 Small trades some of the full attention flexibility standard transformers use for meaningfully faster inference, a specific architectural choice reflecting Microsoft’s broader Phi-3 family emphasis on efficiency-per-parameter over pure benchmark-maximizing design. This tradeoff shows up most clearly in throughput-sensitive deployments where the modest attention flexibility cost is outweighed by real gains in serving cost and latency.

Where Phi-3 Small Fits Between Mini and Medium

Phi-3 Small’s 7B parameter count places it between Phi-3 Mini’s edge-focused 3.8B and Phi-3 Medium’s more capable 14B, giving Microsoft’s lineup a graduated set of options for teams to choose from based on the specific latency, cost, and capability tradeoff a given application demands, without needing to jump directly between the smallest and largest options in the family.

Teams evaluating the Phi-3 family typically benchmark all three sizes against their specific workload before committing, since the right tradeoff point genuinely varies by task and available serving infrastructure.

Historical figures, architectures, and capabilities are for informational purposes only. Not technical, professional, legal, or financial advice. Sources: Benchmark evaluations derived from public developer statements.