In this article we will explore the Microsoft Foundry Local CLI application. We will then write a simple C# program that interacts with a local model that is served by Foundry Local.
Companion Video: https://youtu.be/YD5QKDcb8T8
Prerequisites
Before you begin, ensure you have the following installed on your system:
- .NET 10.0
- Visual Studio Code
What is Microsoft Foundry Local?
Foundry Local is an AI solution that runs entirely on the user's device. It also provides an SDK (C#, JavaScript, Rust, and Python) that helps you build apps that interact with Foundry Local.
Installation
To install Foundry Local, follow these steps:
Windows
winget install Microsoft.FoundryLocal
macOS (only on silicon chips)
brew tap microsoft/foundrylocal
brew trust microsoft/foundrylocal
brew install foundrylocal
Exporing CLI commands
Try the following Foundry Local CLI commands:
Detect the CLI version:
foundry --version
Expected output:
0.8.119
View the list of CLI commands:
foundry --help
Expected output:
Description:
Foundry Local CLI: Run AI models on your device.
🚀 Getting started:
1. To view available models: foundry model list
2. To run a model: foundry model run <model>
EXAMPLES:
foundry model run phi-3-mini-4k
Usage:
foundry [command] [options]
Options:
-?, -h, --help Show help and usage information
--version Show version information
--license Display foundry license information
Commands:
model Discover, run and manage models
cache Manage the local cache
service Manage the local model inference service
Download a model into local cache:
foundry model download qwen2.5-0.5b
Expected output:
Downloading qwen2.5-0.5b-instruct-generic-gpu:4... . . . T R U N C A T E D . . .
[################################## ] 94.85 % [Time remaining: about 1s] [################################## ] 95.22 % [Time remaining: about 1s] [################################## ] 95.59 % [Time remaining: about 1s] [################################## ] 95.96 % [Time remaining: about 1s] [################################## ] 96.33 % [Time remaining: about 1s] [################################## ] 96.69 % [Time remaining: about 1s] [################################## ] 97.06 % [Time remaining: about 1s] [################################### ] 97.43 % [Time remaining: about 1s] [################################### ] 97.79 % [Time remaining: about 1s] [################################### ] 98.16 % [Time remaining: about 1s] [################################### ] 98.53 % [Time remaining: about 1s] [################################### ] 98.90 % [Time remaining: about 1s] [################################### ] 99.27 % [Time remaining: about 1s] [################################### ] 99.63 % [Time remaining: about 1s] [####################################] 100.00 % [Time remaining: about 0s]
- To find model cache location use:
foundry cache location- To find models already downloaded use:
foundry cache ls
List models in local cache:
foundry cache list
Expected output:
💾 qwen2.5-0.5b
qwen2.5-0.5b-instruct-generic-gpu:4
Remove a model from local cache:
foundry cache remove qwen2.5-0.5b
Expected output:
⚠️ This will delete model 'qwen2.5-0.5b (qwen2.5-0.5b-instruct-generic-gpu:4)' from cache.
⚠️ Are you sure you want to delete this model from the local cache? (y/n)
y
Deleted model qwen2.5-0.5b-instruct-generic-gpu:4 from the cache.
Download and run a model:
foundry run qwen2.5-0.5b
Expected output:
Downloading qwen2.5-0.5b-instruct-generic-gpu:4...
[####################################] 100.00 % [Time remaining: about 0s] 105.2 MB/s
🕛 Loading model...
🟢 Model qwen2.5-0.5b-instruct-generic-gpu:4 loaded successfully
Interactive Chat. Enter /? or /help for help.
Press Ctrl+C to cancel generation. Type /exit to leave the chat.
Interactive mode, please enter your prompt
>
If you get an error when running a model, it may mean that your hardware is incompatible with the specific model that was downloaded. You can try looking ar variants of that model for different processor configurations.
Enter /exit to exit Foundry CLI.
foundry model info qwen2.5-0.5b
Expected output:
╭────────────────┬──────────────╮ │ Field │ Value │ ├────────────────┼──────────────┤ │ Alias │ qwen2.5-0.5b │ │ Type │ Chat │ │ Publisher │ Microsoft │ │ License │ apache-2.0 │ │ Capabilities │ tool-calling │ │ Context Length │ 32768 │ │ Tools │ Yes │ ╰────────────────┴──────────────╯ ╭─────────────────┬────────────────┬────────┬────────────────┬────────┬────────╮ │ Variant │ Model ID │ Device │ Execution │ Size │ Cached │ │ │ │ │ Provider │ │ │ ├─────────────────┼────────────────┼────────┼────────────────┼────────┼────────┤ │ qwen2.5-0.5b-in │ qwen2.5-0.5b-i │ GPU │ WebGpuExecutio │ 700 MB │ ● │ │ struct-generic- │ nstruct-generi │ │ nProvider │ │ │ │ gpu │ c-gpu:4 │ │ │ │ │ │ qwen2.5-0.5b-in │ qwen2.5-0.5b-i │ CPU │ CPUExecutionPr │ 822 MB │ ● │ │ struct-generic- │ nstruct-generi │ │ ovider │ │ │ │ cpu │ c-cpu:4 │ │ │ │ │ ╰─────────────────┴────────────────┴────────┴────────────────┴────────┴────────╯
qwen2.5-0.5b-instruct-generic-cpu with: foundry run qwen2.5-0.5b-instruct-generic-cpu
Write a C# console app that converts temperature from C to F and also from F to C.
Expected output:
🧠Thinking...
🤖 Certainly! Below is a simple console application written in C# that takes user input for Celsius temperature and converts it to either Fahrenheit or vice versa.using System; using System.Windows.Forms; namespace TemperatureConverterApp { class Program { static void Main(string[] args) { try { Console.WriteLine("Enter the Temperature (in Celcius): "); int celsius = Convert.ToInt32(Console.ReadLine()); bool convertToFahrenheit = false; string result; if (Convert.ToBoolean(calculateCtoF(celsius)) == true) result = "Celsius to Fahrenheit: "; else result = "Temperature not converted"; result += $"Celsius {result}"; if (convertToFahrenheit) { Console.WriteLine($"Result: Fahrenheits"); return; } Console.WriteLine(result); } catch { Console.WriteLine("Invalid Input"); } } public static bool calculateCtoF(int celsius) { double fahrenheit = 9/5 * celsius + 32; double result = Math.Round(fahrenheit, 1); // Round down to the nearest integer since temperatures must be whole numbers if (fahrenheit.ToString() != "0") return true; else return false; } public static bool calculateFtoc(double fahrenheit) { double celsius = fahrenheit - 32 * 5 / 9; double result = Math.Round(celsius, 1); // Round down to the nearest integer since temperatures must be whole numbers if (celsius.ToString() != "0") return true; else return false; } } }Explanation: 1. Imports: The code starts with importing necessary namespaces (SystemandSystem.Windows.Forms). It's assumed you have this namespace available. 2. Program Class: - This contains the main program logic. 3. calculateCtoF Function: - Calculates the temperature in Fahrenheit given Celsius temperature. - If successful, calculates the equivalent number of degrees Celsius based on the provided value and rounds it to the closest integer. - Outputs the result along with any additional information like "Celsius" or "Temperature not converted". 4. calculateFtoc Function: - Calculates the equivalent temperature in Celcius given the equivalent degrees Fahrenheit. - Same ascalculateCtoF, but used for converting degrees Fahrenheit back to Celsius. Example Usage: If you run the program, pressing the keyboard will take an input to calculate the correct conversion type and output results accordingly. Example Input: 25 - If calculated using Celsius, it will print: "Celsius", then "25". - If converted to Fahrenheit, it will print: "Temperature not converted", then "25". This example only demonstrates basic conversions between the two temperature scales and inputs like non-numerical characters which would throw exceptions during parsing.
Exit CLI mode, type:
/exit
Stop the Foundry service:
foundry server stop
Expected output:
🔴 Service is stopped.
Start the Foundry service:
foundry server start
Expected output:
🟢 Service is Started on http://127.0.0.1:64398/, PID 9459!
Develop a C# app using Foundry Local SDK
The Foundry Local SDK enables you to ship AI features in your applications that are capable of using local AI models through a simple and intuitive API. The SDK abstracts away the complexities of managing AI models and provides a seamless experience for integrating local AI capabilities into your applications.
In the terminal inside a suitable working directory on your computer, create a new console application and add the necessary packages with the following commands:
dotnet new console -o FoundryLocalConsoleApp
cd FoundryLocalConsoleApp
dotnet add package Microsoft.AI.Foundry.Local
dotnet add package Microsoft.Extensions.Logging
Open the project in VS Code by typing this in the same terminal window:
code .
Add this to the .csproj file right above </PropertyGroup>:
<RuntimeIdentifiers>osx-arm64;osx-x64;win-x64;linux-x64</RuntimeIdentifiers>
Replace Program.cs with this code that asks the qwen2.5-0.5b local AI model the
question:
Where did coffee come from?:
1 - Import Namespaces
Purpose:These using statements bring the required libraries into scope.
using Microsoft.AI.Foundry.Local;
using Microsoft.Extensions.Logging;
2 - Create a Cancellation Token
Purpose: Allows long-running operations to be canceled.
CancellationToken ct = new();
3 - Configure Foundry Local
Purpose: Creates runtime settings for Foundry Local.
var config = new Configuration {
AppName = "foundry_local_samples",
LogLevel = Microsoft.AI.Foundry.Local.LogLevel.Information
};
4 - Create a Logger
Purpose: Creates an ILogger.
using var loggerFactory = LoggerFactory.Create(builder => {
// Intentionally no providers configured
});
ILogger logger = loggerFactory.CreateLogger("FoundryLocalConsoleApp");
5 - Initialize Foundry Local
Purpose: Starts the Foundry Local system.
await FoundryLocalManager.CreateAsync(config, logger);
var mgr = FoundryLocalManager.Instance;
6 - Discover Available Execution Providers
Find available hardware acceleration backends.
Examples might include:
- CPU
- DirectML
- CUDA
- ROCm
- ONNX Runtime providers
var eps = mgr.DiscoverEps();
int maxNameLen = 30;
Console.WriteLine("Available execution providers:");
Console.WriteLine($" {"Name".PadRight(maxNameLen)} Registered");
Console.WriteLine($" {new string('─', maxNameLen)} {"──────────"}");
foreach (var ep in eps) {
Console.WriteLine($" {ep.Name.PadRight(maxNameLen)} {ep.IsRegistered}");
}
7 - Download and Register Execution Providers
Purpose: Download and register all execution providers with per-EP progress. EP packages include dependencies and may be large. Download is only required again if a new version of the EP is released. For cross platform builds there is no dynamic EP download and this will return immediately.
Console.WriteLine("\nDownloading execution providers:");
if (eps.Length > 0) {
string currentEp = "";
await mgr.DownloadAndRegisterEpsAsync((epName, percent) => {
if (epName != currentEp) {
if (currentEp != "") {
Console.WriteLine();
}
currentEp = epName;
}
Console.Write($"\r {epName.PadRight(maxNameLen)} {percent,6:F1}%");
});
Console.WriteLine();
} else {
Console.WriteLine("No execution providers to download.");
}
8 - Retrieve the Model Catalog
Purpose: Gets all models available through Foundry Local.
var catalog = await mgr.GetCatalogAsync();
9 - Locate a Model
Purpose: Find a model using an alias.
var model = await catalog.GetModelAsync("qwen2.5-0.5b")
?? throw new Exception("Model not found");
10 - Download the Model
Purpose: Download model files to the local cache.
await model.DownloadAsync(progress => {
Console.Write($"\rDownloading model: {progress:F2}%");
if (progress >= 100f) {
Console.WriteLine();
}
});
11 - Load the Model Into Memory
Purpose: Load the model into memory for inference.
Console.Write($"Loading model {model.Id}...");
await model.LoadAsync();
Console.WriteLine("done.");
12 - Sends a streamed chat request and print final answer.
Purpose: Aski the local AI model a factual question, telling it to answer conservatively if unsure, wait for the full generated response, and then print only the text content of that response to the console.
using (var chatSession = new ChatSession(model)) {
// Enable token streaming so request can produce incremental output.
chatSession.SetStreaming(true);
// Request contains user message and is processed by loaded model.
using var request = new Request();
request.AddItem(MessageItem.User("Where did coffee come from?"));
request.AddItem(MessageItem.System("Say 'I do not know' if you do not know the answer."));
Console.WriteLine("Chat completion response:");
await using var streamingResponse
= chatSession.ProcessStreamingRequestAsync(request, ct);
// Drain stream before reading FinalResponse.
await foreach (var item in streamingResponse) {
Console.Out.Flush();
}
// Dispose final response before session and model are disposed.
using var finalResponse = await streamingResponse.FinalResponse;
foreach (var item in finalResponse) {
if (item is MessageItem message && message.IsSimpleText()) {
Console.Write(message.GetSimpleText());
}
}
}
Console.WriteLine();
13- Clean Up Resources
Purpose: Remove the model from memory.
await model.UnloadAsync();
ℹ️ NOTE - Notice model
qwen2.5-0.5bin the above code (around line 59). Change this to any other model or variant of your choice.
To run the application, type: dotnet run
Expected output:
Available execution providers: Name Registered ────────────────────────────── ────────── WebGpuExecutionProvider False Downloading execution providers: WebGpuExecutionProvider 100.0% Loading model qwen2.5-0.5b-instruct-generic-gpu:4...done. Chat completion response: Coffee originated in the Ethiopian Highlands and later spread to other regions due to various factors such as trade, migration, and disease. It''s believed that people first brought coffee from Ethiopia with them when they migrated to other parts of the world. The earliest known records suggest that the first drink made from coffee was consumed by the Shas people in West Africa. As the demand for coffee grew, more people began to experiment with making their own beverages. By 1750, coffee had been developed into its current form in the Middle East and Asia. It wasn''t until 1826 that an American named Robert Brown invented espresso, which has since become one of the most popular beverages worldwide. The history of coffee is a story of innovation, trade, and cultural exchange over centuries. Coffee has played a significant role in many societies around the world and continues to be enjoyed today through different methods and traditions.
For more information, see the Foundry Local SDK reference.
If you experience an error while running this app, try targeting a specific variant of the model by replacing code:
catalog.GetModelAsync("qwen2.5-0.5b")
with the desired model version, for example:
catalog.GetModelVariantAsync("qwen2.5-0.5b-instruct-generic-gpu:4")
Other commands you can try:
foundry model list
foundry model list --device gpu
foundry model list --search deep
foundry model list --cached
foundry model list --cached --output json
foundry model list --cached --variants
Conclusion
In conclusion, Microsoft Foundry Local offers a powerful and cost-effective solution for developers looking to run and experiment with small language models directly on their own devices. By bypassing the token expenses and network dependencies associated with cloud-hosted platforms like Azure, the platform's intuitive CLI empowers users to efficiently manage model caches, download hardware-tailored model variants, and engage in fast interactive AI prompting entirely offline. Whether you are using it to spin up quick terminal chats or leverage its versatile multi-language SDKs for seamless application integration, Foundry Local effectively removes cloud-based complexities, proving itself to be an invaluable asset for building private, local, and modern AI-driven features.