Showing posts with label Foundry. Show all posts
Showing posts with label Foundry. Show all posts

Friday, July 17, 2026

Coding with Microsoft Foundry Local

In this article we will explore the Microsoft Foundry Local CLI application. We will then write a simple C# program that interacts with a local model that is served by Foundry Local.

Companion Video: https://youtu.be/YD5QKDcb8T8

Prerequisites

Before you begin, ensure you have the following installed on your system:

  • .NET 10.0
  • Visual Studio Code

What is Microsoft Foundry Local?

Foundry Local is an AI solution that runs entirely on the user's device. It also provides an SDK (C#, JavaScript, Rust, and Python) that helps you build apps that interact with Foundry Local.

Installation

To install Foundry Local, follow these steps:

Windows

winget install Microsoft.FoundryLocal

macOS (only on silicon chips)

brew tap microsoft/foundrylocal
brew trust microsoft/foundrylocal
brew install foundrylocal

Exporing CLI commands

Try the following Foundry Local CLI commands:

Detect the CLI version:

foundry --version

Expected output:

0.8.119

View the list of CLI commands:

foundry --help

Expected output:

    Description:

    Foundry Local CLI: Run AI models on your device.
     ðŸš€ Getting started:
     1. To view available models: foundry model list
     2. To run a model: foundry model run <model>

     EXAMPLES:
    
    foundry model run phi-3-mini-4k
    
     Usage:
    
    foundry [command] [options]
     Options:
    -?, -h, --help Show help and usage information
    --version Show version information
    --license Display foundry license information
    
    Commands:
    
    model Discover, run and manage models
    cache Manage the local cache
    service Manage the local model inference service
    

Download a model into local cache:

foundry model download qwen2.5-0.5b

Expected output:

Downloading qwen2.5-0.5b-instruct-generic-gpu:4... 

 . . . T R U N C A T E D . . .  

[################################## ] 94.85 % [Time remaining: about 1s] [################################## ] 95.22 % [Time remaining: about 1s] [################################## ] 95.59 % [Time remaining: about 1s] [################################## ] 95.96 % [Time remaining: about 1s] [################################## ] 96.33 % [Time remaining: about 1s] [################################## ] 96.69 % [Time remaining: about 1s] [################################## ] 97.06 % [Time remaining: about 1s] [################################### ] 97.43 % [Time remaining: about 1s] [################################### ] 97.79 % [Time remaining: about 1s] [################################### ] 98.16 % [Time remaining: about 1s] [################################### ] 98.53 % [Time remaining: about 1s] [################################### ] 98.90 % [Time remaining: about 1s] [################################### ] 99.27 % [Time remaining: about 1s] [################################### ] 99.63 % [Time remaining: about 1s] [####################################] 100.00 % [Time remaining: about 0s]

💡 Tip:

- To find model cache location use: foundry cache location
- To find models already downloaded use: foundry cache ls


List models in local cache:

foundry cache list

Expected output:

    💾 qwen2.5-0.5b             
    qwen2.5-0.5b-instruct-generic-gpu:4

Remove a model from local cache:

foundry cache remove qwen2.5-0.5b

Expected output:

    ⚠️ This will delete model 'qwen2.5-0.5b (qwen2.5-0.5b-instruct-generic-gpu:4)' from cache.
    ⚠️ Are you sure you want to delete this model from the local cache? (y/n)
    y
    Deleted model qwen2.5-0.5b-instruct-generic-gpu:4 from the cache.

Download and run a model:

foundry run qwen2.5-0.5b

Expected output:

Downloading qwen2.5-0.5b-instruct-generic-gpu:4...
[####################################] 100.00 % [Time remaining: about 0s]  105.2 MB/s
🕛 Loading model...
🟢 Model qwen2.5-0.5b-instruct-generic-gpu:4 loaded successfully

Interactive Chat. Enter /? or /help for help.
Press Ctrl+C to cancel generation. Type /exit to leave the chat.

Interactive mode, please enter your prompt
>

💡 TIP

If you get an error when running a model, it may mean that your hardware is incompatible with the specific model that was downloaded. You can try looking ar variants of that model for different processor configurations. 

Enter /exit to exit Foundry CLI.

To view the variants for a model (Example: qwen2.5-0.5b), type the following command: 

foundry model info qwen2.5-0.5b

Expected output:

╭────────────────┬──────────────╮
│ Field          │ Value        │
├────────────────┼──────────────┤
│ Alias          │ qwen2.5-0.5b │
│ Type           │ Chat         │
│ Publisher      │ Microsoft    │
│ License        │ apache-2.0   │
│ Capabilities   │ tool-calling │
│ Context Length │ 32768        │
│ Tools          │ Yes          │
╰────────────────┴──────────────╯
╭─────────────────┬────────────────┬────────┬────────────────┬────────┬────────╮
│ Variant         │ Model ID       │ Device │ Execution      │ Size   │ Cached │
│                 │                │        │ Provider       │        │        │
├─────────────────┼────────────────┼────────┼────────────────┼────────┼────────┤
│ qwen2.5-0.5b-in │ qwen2.5-0.5b-i │ GPU    │ WebGpuExecutio │ 700 MB │ ●      │
│ struct-generic- │ nstruct-generi │        │ nProvider      │        │        │
│ gpu             │ c-gpu:4        │        │                │        │        │
│ qwen2.5-0.5b-in │ qwen2.5-0.5b-i │ CPU    │ CPUExecutionPr │ 822 MB │ ●      │
│ struct-generic- │ nstruct-generi │        │ ovider         │        │        │
│ cpu             │ c-cpu:4        │        │                │        │        │
╰─────────────────┴────────────────┴────────┴────────────────┴────────┴────────╯

You can then try one of the variants like qwen2.5-0.5b-instruct-generic-cpu with: 
foundry run qwen2.5-0.5b-instruct-generic-cpu

The above puts the user in a CLI interactive mode. You can then send prompt to the model. Let's ask the model to create for us a C# program that does temperature conversion. Type in the following prompt:

Write a C# console app that converts temperature from C to F and also from F to C.

Expected output:

    🧠 Thinking...
🤖 Certainly!  Below is a simple console application written in C# that takes user input for Celsius temperature and converts it to either Fahrenheit or vice versa. using System; using System.Windows.Forms; namespace TemperatureConverterApp { class Program { static void Main(string[] args) { try { Console.WriteLine("Enter the Temperature (in Celcius): "); int celsius = Convert.ToInt32(Console.ReadLine()); bool convertToFahrenheit = false; string result; if (Convert.ToBoolean(calculateCtoF(celsius)) == true) result = "Celsius to Fahrenheit: "; else result = "Temperature not converted"; result += $"Celsius {result}"; if (convertToFahrenheit) { Console.WriteLine($"Result: Fahrenheits"); return; } Console.WriteLine(result); } catch { Console.WriteLine("Invalid Input"); } } public static bool calculateCtoF(int celsius) { double fahrenheit = 9/5 * celsius + 32; double result = Math.Round(fahrenheit, 1); // Round down to the nearest integer since temperatures must be whole numbers if (fahrenheit.ToString() != "0") return true; else return false; } public static bool calculateFtoc(double fahrenheit) { double celsius = fahrenheit - 32 * 5 / 9; double result = Math.Round(celsius, 1); // Round down to the nearest integer since temperatures must be whole numbers if (celsius.ToString() != "0") return true; else return false; } } } Explanation:  1. Imports: The code starts with importing necessary namespaces (System and System.Windows.Forms). It's assumed you have this namespace available.  2. Program Class:      - This contains the main program logic.  3. calculateCtoF Function:      - Calculates the temperature in Fahrenheit given Celsius temperature.      - If successful, calculates the equivalent number of degrees Celsius based on the provided value and rounds it to the closest integer.       - Outputs the result along with any additional information like "Celsius" or "Temperature not converted".  4. calculateFtoc Function:      - Calculates the equivalent temperature in Celcius given the equivalent degrees Fahrenheit.      - Same as calculateCtoF, but used for converting degrees Fahrenheit back to Celsius.  Example Usage:  If you run the program, pressing the keyboard will take an input to calculate the correct conversion type and output results accordingly.  Example Input: 25      - If calculated using Celsius, it will print: "Celsius", then "25".      - If converted to Fahrenheit, it will print: "Temperature not converted", then "25".  This example only demonstrates basic conversions between the two temperature scales and inputs like non-numerical characters which would throw exceptions during parsing.

Exit CLI mode, type:

/exit

Stop the Foundry service:

foundry server stop

Expected output:

    🔴 Service is stopped.
    

Start the Foundry service:

foundry server start

Expected output:

    🟢 Service is Started on http://127.0.0.1:64398/, PID 9459!
    

Develop a C# app using Foundry Local SDK

The Foundry Local SDK enables you to ship AI features in your applications that are capable of using local AI models through a simple and intuitive API. The SDK abstracts away the complexities of managing AI models and provides a seamless experience for integrating local AI capabilities into your applications.

In the terminal inside a suitable working directory on your computer, create a new console application and add the necessary packages with the following commands:

dotnet new console -o FoundryLocalConsoleApp
cd FoundryLocalConsoleApp
dotnet add package Microsoft.AI.Foundry.Local
dotnet add package Microsoft.Extensions.Logging

Open the project in VS Code by typing this in the same terminal window:

code .

Add this to the .csproj file right above </PropertyGroup>:

<RuntimeIdentifiers>osx-arm64;osx-x64;win-x64;linux-x64</RuntimeIdentifiers>

Replace Program.cs with this code that asks the qwen2.5-0.5b local AI model the question: Where did coffee come from?:

1 - Import Namespaces

Purpose:These using statements bring the required libraries into scope.

using Microsoft.AI.Foundry.Local;
using Microsoft.Extensions.Logging;
2 - Create a Cancellation Token

Purpose: Allows long-running operations to be canceled.

CancellationToken ct = new();
3 - Configure Foundry Local

Purpose: Creates runtime settings for Foundry Local.

var config = new Configuration {
   AppName = "foundry_local_samples",
   LogLevel = Microsoft.AI.Foundry.Local.LogLevel.Information
};
4 - Create a Logger

Purpose: Creates an ILogger.

using var loggerFactory = LoggerFactory.Create(builder => {
   // Intentionally no providers configured
});

ILogger logger = loggerFactory.CreateLogger("FoundryLocalConsoleApp");
5 - Initialize Foundry Local

Purpose: Starts the Foundry Local system.

await FoundryLocalManager.CreateAsync(config, logger);
var mgr = FoundryLocalManager.Instance;
6 - Discover Available Execution Providers

Find available hardware acceleration backends.

Examples might include:

  • CPU
  • DirectML
  • CUDA
  • ROCm
  • ONNX Runtime providers
var eps = mgr.DiscoverEps();
int maxNameLen = 30;
Console.WriteLine("Available execution providers:");
Console.WriteLine($" {"Name".PadRight(maxNameLen)} Registered");
Console.WriteLine($" {new string('─', maxNameLen)} {"──────────"}");
foreach (var ep in eps) {
    Console.WriteLine($" {ep.Name.PadRight(maxNameLen)} {ep.IsRegistered}");
}
7 - Download and Register Execution Providers

Purpose: Download and register all execution providers with per-EP progress. EP packages include dependencies and may be large. Download is only required again if a new version of the EP is released. For cross platform builds there is no dynamic EP download and this will return immediately.

Console.WriteLine("\nDownloading execution providers:");
if (eps.Length > 0) {
   string currentEp = "";
   await mgr.DownloadAndRegisterEpsAsync((epName, percent) => {
      if (epName != currentEp) {
         if (currentEp != "") {
            Console.WriteLine();
         }
         currentEp = epName;
      }
      Console.Write($"\r {epName.PadRight(maxNameLen)} {percent,6:F1}%");
   });
   Console.WriteLine();
} else {
   Console.WriteLine("No execution providers to download.");
}
8 - Retrieve the Model Catalog

Purpose: Gets all models available through Foundry Local.

var catalog = await mgr.GetCatalogAsync();
9 - Locate a Model

Purpose: Find a model using an alias.

var model = await catalog.GetModelAsync("qwen2.5-0.5b")
    ?? throw new Exception("Model not found");
10 - Download the Model

Purpose: Download model files to the local cache.

await model.DownloadAsync(progress => {
   Console.Write($"\rDownloading model: {progress:F2}%");
   if (progress >= 100f) {
      Console.WriteLine();
   }
});
11 - Load the Model Into Memory

Purpose: Load the model into memory for inference.

Console.Write($"Loading model {model.Id}...");
await model.LoadAsync();
Console.WriteLine("done.");
12 - Sends a streamed chat request and print final answer.

Purpose: Aski the local AI model a factual question, telling it to answer conservatively if unsure, wait for the full generated response, and then print only the text content of that response to the console.

using (var chatSession = new ChatSession(model)) {
   // Enable token streaming so request can produce incremental output.
   chatSession.SetStreaming(true);

   // Request contains user message and is processed by loaded model.
   using var request = new Request();
   request.AddItem(MessageItem.User("Where did coffee come from?"));
   request.AddItem(MessageItem.System("Say 'I do not know' if you do not know the answer."));

   Console.WriteLine("Chat completion response:");
   await using var streamingResponse
      = chatSession.ProcessStreamingRequestAsync(request, ct);

   // Drain stream before reading FinalResponse.
   await foreach (var item in streamingResponse) {
      Console.Out.Flush();
   }

   // Dispose final response before session and model are disposed.
   using var finalResponse = await streamingResponse.FinalResponse;
   foreach (var item in finalResponse) {
      if (item is MessageItem message && message.IsSimpleText()) {
         Console.Write(message.GetSimpleText());
      }
   }
}
Console.WriteLine();
13- Clean Up Resources

Purpose: Remove the model from memory.

await model.UnloadAsync();

ℹ️ NOTE - Notice model qwen2.5-0.5b in the above code (around line 59). Change this to any other model or variant of your choice.

To run the application, type: dotnet run

Expected output:

  Available execution providers:
  Name                            Registered
  ──────────────────────────────  ──────────
  WebGpuExecutionProvider         False

Downloading execution providers:
  WebGpuExecutionProvider          100.0%
Loading model qwen2.5-0.5b-instruct-generic-gpu:4...done.
Chat completion response:

Coffee originated in the Ethiopian Highlands and later spread to other regions 
due to various factors such as trade, migration, and disease. It''s believed that 
people first brought coffee from Ethiopia with them when they migrated to other 
parts of the world. The earliest known records suggest that the first drink made 
from coffee was consumed by the Shas people in West Africa. As the demand for coffee grew, 
more people began to experiment with making their own beverages. By 1750, coffee had 
been developed into its current form in the Middle East and Asia. It wasn''t until 1826 
that an American named Robert Brown invented espresso, which has since become one of the 
most popular beverages worldwide. The history of coffee is a story of innovation, trade, 
and cultural exchange over centuries. Coffee has played a significant role in many societies 
around the world and continues to be enjoyed today through different methods and traditions.

For more information, see the Foundry Local SDK reference.

💡 Tip:

If you experience an error while running this app, try targeting a specific variant of the model by replacing code:

catalog.GetModelAsync("qwen2.5-0.5b")

with the desired model version, for example:

catalog.GetModelVariantAsync("qwen2.5-0.5b-instruct-generic-gpu:4")

Other commands you can try:

foundry model list
foundry model list --device gpu
foundry model list --search deep
foundry model list --cached
foundry model list --cached --output json
foundry model list --cached --variants
    

Conclusion

In conclusion, Microsoft Foundry Local offers a powerful and cost-effective solution for developers looking to run and experiment with small language models directly on their own devices. By bypassing the token expenses and network dependencies associated with cloud-hosted platforms like Azure, the platform's intuitive CLI empowers users to efficiently manage model caches, download hardware-tailored model variants, and engage in fast interactive AI prompting entirely offline. Whether you are using it to spin up quick terminal chats or leverage its versatile multi-language SDKs for seamless application integration, Foundry Local effectively removes cloud-based complexities, proving itself to be an invaluable asset for building private, local, and modern AI-driven features.