This configuration dialog contains all settings that control IMatch AI, semantic search, and related features.
This is an experimental feature. This settings page is visible only when experimental features are enabled.
Semantic search uses AI to generate so-called embeddings from selected tag values. See Semantic Search with AI for detailed information.
This is an experimental feature. You must explicitly enable the IMatch AI semantic search engine in Edit menu > Preferences > Application: Experimental Features. This settings group is available only when experimental features are enabled in the same dialog.
The following settings control how IMatch creates embeddings for semantic search.
Model | Select the model for which you want to configure tags and options. IMatch may offer one or more models in this drop-down list, depending on which models you have downloaded. See Downloading Models below for more information about supported models. |
Tags | This list controls which tags IMatch uses to generate embeddings for this model. Useful tags typically include description and hierarchical keywords, if they contain data. The best results are achieved when your image descriptions are fairly detailed, or when you maintain a separate tag that contains text optimized for generating embeddings. See Semantic Search with AI for detailed information. The information in the tags is crucial for the quality of semantic search. Semantic search works very well when the source data is good. |
Minimum Length | The minimum length of a tag value to consider for embedding. Avoid using very short texts or phrases, as they can negatively affect embedding and search quality. The default is 5, which is a compromise in case you create embeddings from keywords. If you create embeddings from descriptions or dedicated embedding tags, you can increase this value to 10 or more to keep empty or incomplete descriptions out of the embedding index. Junk data can pollute the index, making the search less efficient. |
Ignore Files in this category | This feature is mostly meant for testing purposes. While you experiment with embeddings to understand which tags to use and which AutoTagger prompts to use to fill dedicated embedding tags, and so on, it makes sense to limit the number of files for which you create embeddings. A typical use for this is to create a new category named Embedding_Test_Set and assign a mix of 500 to 1,000 images to this category. This is your test set. Now create a category named Embedding_Ignore with the following formula: "@All" NOT "Embedding_Test_Set" Then select this category for the Ignore Files in this category option. The embedding engine in IMatch now considers only the files in the Embedding_Test_Set category and ignores all other files in the database. If your test set is in a folder instead of a category, you can change the formula to something like this: "@All" NOT "@RFolder[c:\images\test_set]" Once your tests are complete and you are satisfied with your setup, remove the ignore category by unchecking it in the Category Selector dialog. Then run a database diagnosis to schedule all files with missing embeddings for background processing. Using @AllAnother way to use this feature is to set When you use this option, IMatch will not automatically recreate embeddings when you change the underlying tag because all files are excluded from embeddings. You have to select the files you have changed, open this dialog and then use the the Recreate All button in combination with the Only for the n files selected option to recreate the embeddings yourself. |
When you switch to another model or you close the settings dialog with OK, IMatch analyzes the changes you have made and applies them to the database.
Tags added | If you add tags, IMatch adds tasks to the background processing queue to generate embeddings for the added tags for all non-excluded files. |
Tags removed | IMatch removes embedding data for the tags you have removed. |
Minimum length changed | IMatch recreates all embeddings for all non-excluded files to apply the new length settings. |
Since applying such changes to large sets of files can be time-consuming, IMatch asks if you really want to apply your changes. Only when you confirm are the changes applied and background tasks scheduled.
When you change the contents of a tag referenced by one or more models, IMatch AI automatically updates the corresponding embeddings in the background.
Use the indicators in the Info & Activity Panel or the Dashboard to see which tasks are running in the background.
The database diagnosis checks for embeddings linked to models that are no longer available and removes them. It also checks for missing embeddings and adds tasks to the background processing queue to recreate them when needed. This can happen when you use the options in this dialog to create embeddings only for a selection of files.
To use semantic search and create embeddings with IMatch AI, you have to download at least one of the supported models.
Although embedding models are usually much smaller than the AI models that can run locally in Ollama or LM Studio for AutoTagger, they are still too large to include in the regular IMatch installer. In addition, not all IMatch users will use embeddings, so it makes no sense to force them to download models with the IMatch installer.
Download models to the folder configured for the Embedding Model Folder option shown in this dialog.
This list may change over time. Let us know about interesting models that we should add via the IMatch user community.
After downloading a new model, you must restart IMatch so the new models are detected.
Google Gemma Embedding | This is a small but powerful 300M parameter model that runs well even on older and slower hardware. It has an 8K token context length and supports text in 100 languages. It needs only about 600 MB of graphics card memory (VRAM) per thread for a full offload to the GPU. This means it can run entirely on the graphics card, even on older cards or cards with only one or two GB of VRAM. This is the fastest embedding model currently supported by IMatch, with the smallest memory footprint. It should work well for the majority of IMatch users. Try this model first. Download the model using this link and store it in the embedding model folder shown in this dialog. If your graphics card has 2 GB of VRAM, try 2 threads. If it has 4 GB of VRAM, try 6 threads. More threads means more embeddings created in parallel, which is a lot faster. |
The models below are probably overkill for most use cases in IMatch. | |
Qwen3-Embedding-4B Q6 | This capable 4B-parameter embedding model supports text in 100 languages. Use it for the best embedding quality if your system is equipped with a supported, fast graphics card with at least 10 GB of VRAM. It has a 32K token context window which enables it to support very long texts for encoding ‐ which is not a typical use case for IMatch. Download the model using this link and store it in the embedding model folder shown in this dialog. Set the threads option to 1 for this model. If your PC has a powerful graphics card with 16 GB of VRAM and 32 GB of RAM, use 2 threads. This model requires a lot of VRAM and RAM when generating embeddings. |
Qwen3-Embedding-4B Q8 | This is a large 8B-parameter model and the most capable embedding model currently supported. It supports text in 100 languages. Use it for the best embedding quality if your system is equipped with a supported, fast graphics card with at least 16 GB of VRAM. It has a 32K token context window which enables it to support very long texts for encoding ‐ which is not a typical use case for IMatch. Download the model using this link and store it in the embedding model folder shown in this dialog. Set the threads option to 1 for this model. If your PC has a powerful graphics card with at least 20 GB of VRAM and 32 GB of RAM, use 2 threads. This model requires a lot of VRAM and RAM when generating embeddings. |
IMatch AI, introduced in IMatch 2026, is based on the renowned and popular llama.cpp open-source project, like many other AI tools (including Ollama and Unsloth Desktop).
Currently the IMatch AI is used for generating embeddings for semantic search. We plan to use the integrated AI for other features in the future.
IMatch AI can run on a wide range of hardware configurations. The minimum requirement is that the AVX processor feature be available. IMatch has required AVX since we introduced face recognition years ago.
If the processor in your PC supports AVX2, the IMatch AI can generate embeddings 5 to 10 times faster than with AVX alone.
If your PC has one of the many graphics cards supported by the Vulkan Graphics Library that llama.cpp uses, IMatch AI automatically offloads work to your graphics card for much faster processing, according to the settings you configure in this dialog.
Open the IMatch Dashboard and look at the Info section. It shows whether your processor has AVX2 support and whether one or more compatible graphics cards were detected:

The system used for the screenshot has both a built-in, low-end graphics processor and a separate NVIDIA graphics card. The processor supports AVX2.
IMatch AI works on the hardware IMatch users have. It is just slower or faster depending on the system. Processors with AVX2 are 5 to 10 times faster for AI work than processors with AVX only. A modern graphics card with 8 or more GB VRAM can be 5 to 50 times faster than AVX2.
If your PC is older and has only AVX, working with local AI, including the IMatch AI, can be very slow and frustrating.
There is nothing we can do about this; sorry.
It is impossible to automatically adapt all IMatch AI performance features to the hardware users have or the models they run. This is why IMatch ships with safe defaults (CPU-only, 1 thread) and allows the user to fine-tune these settings with the controls explained below.
The settings are stored and applied per model. Different models have different demands for VRAM, RAM, and CPU resources. A small model such as Gemma Embedding may run fine with 4 or even 8 threads on your system, but for the large Qwen 4B or 8B models, only 1 or 2 threads may be feasible, because these models requires multiple GB of VRAM and regular RAM per thread.
This setting controls how many parallel threads IMatch can use when processing AI requests. The default is 1, which is a safe default. It gets the work done, but may leave resources unused and take longer.
For CPU-only operation, set the number of threads to 0, or set it to the number of threads reported for your CPU in Windows Task Manager or a lower value.
Setting the number of threads to 0 selects auto. IMatch then tries to optimize the setting for your system automatically, respecting the performance profile you have configured.
If you notice that the IMatch AI uses too many CPU or GPU (graphics card) resources, set the number of threads to a small value, such as 1 or 2, and test again. Then increase the value until you find a good balance between system utilization and AI performance.
If you run larger AI models (e.g., Qwen 4B and especially Qwen 8B instead of Gemma Embedding), use one or two threads only. Each thread requires multiple GB of VRAM and RAM with these models.
While you experiment with threads, keep an eye on the Performance tab in Windows Task Manager. Your goal is to use as many threads as possible without maxing out RAM, VRAM, or processor resources.
This setting controls how many AI model layers IMatch AI is allowed to move to the GPU (graphics card) in your system. The more layers it can move to the graphics card instead of keeping them in normal (slow) RAM, the faster the AI will run.
By default, IMatch uses the value 0, which means CPU-only. This is the safest but slowest setting.
The value 100 means "offload all layers", which is the fastest and preferred option. The IMatch AI tries to offload as many layers as possible. If this causes issues with your graphics card, even crashes, configure a smaller value of 5 or 10 and re-test.
The behavior often depends on the graphics card, its drivers, other software you run that uses graphics card resources, and the llama.cpp software IMatch uses to run AI locally, so it may not always be possible to predict. This is why IMatch gives you options to control how much graphics card memory (VRAM) can be used for offloading.
This setting controls how aggressively the IMatch AI can use graphics card resources.
Auto | Recommended for most people. IMatch picks a safe default based on your GPU. |
Minimal | Safest choice for weak hardware, CPU-only use, or if you hit memory issues. |
Low | Good for modest laptops or smaller GPUs when you want stability first. |
Balanced | Best general-purpose manual choice for modern GPUs. |
High | For strong GPUs and users who want maximum throughput, especially on larger batches. |
This is the folder where IMatch looks for the model.json file and the models you have downloaded. By default this folder is %APPDATA%\photools.com\IMatch6\AI\Embeddings.
This is the folder into which you download models.
If you run out of disk space on drive C:, you can move the complete Embeddings folder to another disk and select the new folder for this option. Restart IMatch after changing the folder.
The advanced options for embeddings may change or be removed once semantic search leaves experimental status. They are mostly intended for testing purposes.
This button schedules all non-excluded files to have their embeddings recreated. In combination with the Only for the n files selected option, you can recreate embeddings for the files selected in the active File Window only.
This is helpful when you experiment with different models, performance settings, and so on. Recreate embeddings with the new settings for a small test set of files selected in the active File Window.
If you recreate embeddings for selected files, the exclusion category is ignored and all selected files are processed.
This button assigns all files with at least one embedding for the currently selected model to the selected category. This is helpful for testing purposes or when you have lost track of files while working with exclusion categories.
IMatch uses an external helper process named IMatchLLHHelper.exe. The IMatch AI starts and terminates these processes as needed, running multiple processes simultaneously, depending on the maximum number of threads allowed in the settings and the current system load.
The reason for using helper processes is isolation. Sometimes, llama.cpp encounters problems and calls abort to terminate the process running it. Without isolation, this would also terminate IMatch—which would obviously be bad.
llama.cpp can terminate the helper process without affecting IMatch. IMatch recognizes when a helper process has been terminated, records this in the log file, and starts a new helper process when needed.
Because the helper processes run independently of IMatch, they cannot access the IMatch log file. Instead, each helper process creates its own log file using the naming scheme IMATCH6_LLMH_LOG-nn.txt, where nn is a sequential number. The more threads your system can run, the more log files are created in the temporary folder.
IMatch reuses the log file names during each session. If you want to analyze an issue related to the IMatch AI, make copies of the log files before restarting IMatch.
The helper process log files are useful when a model does not work correctly, when your GPU is not being used, or when IMatch AI falls back to CPU-only mode.