Home APIHow to Query and Compare Multiple LLMs on Linux

How to Query and Compare Multiple LLMs on Linux

By sk
3 views 15 mins read

Choosing a suitable large language model (LLM) for a task is hard to do well without running the same prompt through several of them. Each provider has its own account, API key and often its own SDK, so putting Claude, DeepSeek and an open-weights model side by side takes a lot of setup.

In this tutorial, we are going to connect five different models through a single OpenAI-compatible endpoint and run a short Python script that sends the same prompt to each of them, allowing us to compare their responses, response times, and token usage.

Everything here tested on a Linux machine. The install commands below use apt for Debian and Ubuntu. On Fedora and RHEL, use dnf instead.

One Endpoint, Many Models

For the purpose of this guide, we will be using AI/ML API, an API aggregator which serves models from OpenAI, Anthropic, DeepSeek, xAI, Z.ai and other developers through one OpenAI-compatible endpoint at https://api.aimlapi.com/v1.

Since it uses the OpenAI protocol, you can install the standard OpenAI Python SDK, point base_url at this endpoint and switch models by changing one string.

The 5 models used in this tutorial:

ModelDeveloperModel IDPrice per 1M tokens (input / output)
Claude Sonnet 5Anthropicanthropic/claude-sonnet-5$2.60 / $13.00
DeepSeek V4 ProDeepSeekdeepseek/deepseek-v4-pro$0.57 / $1.13
DeepSeek V4 FlashDeepSeekdeepseek/deepseek-v4-flash$0.39 / $1.56
Grok Build 0.1xAIx-ai/grok-build-0-1$1.30 / $2.60
GLM 5.2Z.aizhipu/glm-5-2$1.82 / $5.72

Prices are from the AI/ML API model pages as of 24 September 2026.

⚠️ What to Know Before Paying for API Aggregators

API aggregators can be a convenient way to test multiple models through a single interface, but exercise caution before using them in production or making significant financial commitments.

Start with the smallest deposit available, and disable auto-top-up until you're confident in the service. Test latency and reliability with your actual workload, and carefully verify token usage against your billing.

Important: As of September 2026, AI/ML API's documentation explicitly states that top-ups are non-refundable, with available top-ups starting at $20. Its Terms and Conditions clearly state that all purchases are final and non-refundable.

We strongly recommend you to review their Terms and Conditions, Account & Billing, and Plans pages before adding funds.

Now let us go ahead to compare multiple LLMs via AI/ML API on Linux.

Prerequisites

You need Python 3.9 or newer, pip and an API key. On Debian or Ubuntu, run the following command to install them:

sudo apt install python3 python3-venv python3-pip

Create a key in your AI/ML API dashboard at API keys page.

Billing is pay-as-you-go, and the dashboard has a web Playground where you can try a model before scripting against it.

Step 1: Set Up an Isolated Environment

Install the SDK inside a virtual environment, not system-wide:

mkdir llm-compare && cd llm-compare  
python3 -m venv .venv
source .venv/bin/activate
pip install openai

Then export the key as an environment variable so it is not hard-coded in the script:

export AIMLAPI_KEY="your_api_key_here"

To keep the key out of your shell history, add this line to ~/.bashrc with a text editor instead of typing it at the prompt.

Step 2: Check the Endpoint with Curl

Before writing any Python script, confirm that the endpoint answers. It takes a standard POST request to /v1/chat/completions with a Bearer token:

curl https://api.aimlapi.com/v1/chat/completions \
-H "Authorization: Bearer $AIMLAPI_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek/deepseek-v4-flash",
"messages": [{"role": "user", "content": "Reply with the single word: ok"}]
}'

The response is the usual OpenAI JSON, with the answer in choices[0].message.content and token counts in usage. If this request works, the script below will work too.

Step 3: Python Script for Comparing Multiple LLMs

This script sends the same prompt to every model in a list, measures how long each request takes and prints the answers one after another. Save it as compare_llms.py:

#!/usr/bin/env python3
"""compare_llms.py - send one prompt to several LLMs through a single
OpenAI-compatible endpoint (AI/ML API) and compare answers, latency and tokens."""

import os
import sys
import time

from openai import OpenAI

BASE_URL = os.environ.get("OPENAI_BASE_URL", "https://api.aimlapi.com/v1")
API_KEY = os.environ.get("AIMLAPI_KEY") or os.environ.get("OPENAI_API_KEY")

# Edit this list, or override it with the MODELS env var (comma-separated).
MODELS = os.environ.get(
"MODELS",
"anthropic/claude-sonnet-5,"
"deepseek/deepseek-v4-pro,"
"deepseek/deepseek-v4-flash,"
"x-ai/grok-build-0-1,"
"zhipu/glm-5-2",
).split(",")


def main() -> int:
if not API_KEY:
sys.exit("Set AIMLAPI_KEY first (create a key at https://aimlapi.com/app/keys).")

prompt = " ".join(sys.argv[1:]) or "Explain what a Linux inode is, in two sentences."
client = OpenAI(base_url=BASE_URL, api_key=API_KEY)

print(f"Prompt: {prompt}\n" + "=" * 60)

for raw in MODELS:
model = raw.strip()

if not model:
continue

start = time.perf_counter()

try:
resp = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": prompt}],
)
except Exception as exc: # network, auth, unknown model, etc.
print(f"\n### {model}\n ERROR: {exc}")
continue

elapsed = time.perf_counter() - start
answer = (resp.choices[0].message.content or "").strip()
usage = resp.usage
tokens = f"{usage.prompt_tokens} in / {usage.completion_tokens} out" if usage else "n/a"

print(f"\n### {model} ({elapsed:.2f}s, {tokens})\n{answer}")

return 0


if __name__ == "__main__":
raise SystemExit(main())

Some notes on the design:

  • The model list comes from the MODELS environment variable and falls back to the 5 models above. You can edit the file or run MODELS="deepseek/deepseek-v4-flash,zhipu/glm-5-2" python3 compare_llms.py without touching the code.
  • Each request is wrapped in try/except. A wrong model ID or a network error prints an error for that model, and the script moves on to the next one.
  • The script prints input and output token counts from the usage field, so you can multiply them by the prices in the table above and compare the cost of the same answer across models.
  • It does not set temperature. DeepSeek V4 Pro and Claude Sonnet 5 have thinking modes, and some reasoning configurations ignore or reject sampling parameters. If you want steadier answers from non-reasoning models, add temperature=0.3 to the create() call.

Step 4: Query and Compare Multiple LLMs

Pass your prompt as arguments, or run the script without arguments to use the default prompt:

python3 compare_llms.py "Explain what a Linux inode is, in two sentences."

The output lists the models one after another, each tagged with the response time and the token count. This is the output from our run on 24 September 2026:

Prompt: Explain what a Linux inode is, in two sentences.  
============================================================

### anthropic/claude-sonnet-5 (8.44s, 24 in / 126 out)
An inode (index node) is a data structure on a filesystem that stores metadata about a file or directory—such as its size, permissions, ownership, timestamps, and pointers to the actual data blocks on disk—but notably does not store the file's name. Each file is linked to a unique inode via directory entries, allowing multiple filenames (hard links) to reference the same underlying data through a shared inode.

### deepseek/deepseek-v4-pro (4.65s, 95 in / 149 out)
A Linux inode is a filesystem data structure that stores metadata about a file or directory, such as permissions, ownership, timestamps, and pointers to its data blocks. It does not store the file’s name or its actual content; instead, each inode is identified by a unique inode number within its filesystem.

### deepseek/deepseek-v4-flash (2.51s, 42 in / 279 out)
An inode is a filesystem data structure that stores a file’s metadata—such as its type, permissions, ownership, timestamps, size, and pointers to data blocks—but not its name or actual data content. Every file has an inode number unique within its filesystem, and directory entries map human-readable filenames to those inode numbers, which is why hard links can share one inode.

### x-ai/grok-build-0-1 (3.92s, 197 in / 323 out)
In Linux, an inode is a data structure that stores metadata about a file or directory, such as its permissions, ownership, timestamps, size, and pointers to the data blocks on disk. It does not contain the file's name or its actual content; the filename is stored in the directory entry, which references the inode number.

### zhipu/glm-5-2 (15.70s, 24 in / 709 out)
A Linux inode is a data structure that stores metadata about a file or directory, such as its permissions, owner, size, and physical location on the disk. It acts as the bridge between a file's human-readable name and its actual data blocks, notably excluding the file name itself, which is instead stored in the directory structure.

Your times and token counts will differ from run to run, but a few things stand out. All 5 models gave a correct two-sentence answer. DeepSeek V4 Flash was the fastest at 2.51 seconds, and GLM 5.2 the slowest at 15.70 seconds.

The token counts are the more useful part. The prompt was the same for every model, yet input counts range from 24 to 197 tokens, because each model uses its own tokenizer and some providers can add their own instructions to the request. Output counts vary even more. GLM 5.2 billed 709 output tokens for an answer of about 60 words, most likely because it is a reasoning model and the tokens it spends thinking are billed as output.

Here is what this one request cost on each model, using the prices from the table above:

ModelTokens (in / out)Cost of this request
DeepSeek V4 Pro95 / 149$0.00022
DeepSeek V4 Flash42 / 279$0.00045
Grok Build 0.1197 / 323$0.0011
Claude Sonnet 524 / 126$0.0017
GLM 5.224 / 709$0.0041

DeepSeek V4 Flash has the lowest input price of these models. On this prompt, though, it cost twice as much as DeepSeek V4 Pro, because it produced more output tokens and its output price is higher. Price lists alone are not enough to compare models, so run your own prompts and check the actual token counts.

Replace the default prompt with your own work. For example, ask the models to explain a bug, review a shell script or generate a config file.

Step 5: What to Compare, and How to Pick Models

Three things usually decide which model to use.

  1. Latency. The script prints it for every request. Run it several times, because the first call to a model is often slower than the rest.
  2. Cost. Each model has its own price per million tokens, and for long answers the output price matters most. If you send thousands of requests a day, pick the cheapest model that still answers your task well enough. The savings add up quickly at that volume.
  3. Quality. Only you can judge it, and seeing the answers side by side makes that easier.

Here is what sets each model apart.

  • Claude Sonnet 5 is Anthropic's model for coding and agent workflows, released on 30 June 2026. Anthropic says it comes close to its top Opus model at a lower price. In our run it gave the most detailed answer with the fewest output tokens, but its price per token is the highest in this list.
  • DeepSeek V4 Pro is the largest model here, with 1.6T parameters. Since August 2026 you can set its reasoning effort to low, high or max, so one model covers both quick answers and harder tasks. It was also the cheapest request in our run.
  • DeepSeek V4 Flash is built for speed. It activates only 13B of its 284B parameters per request and was the fastest model in our test at 2.51 seconds. It suits high-volume tasks where a quick answer matters more than depth.
  • Grok Build 0.1 is the only model in this list trained specifically for coding agents, with support for MCP tools. xAI released it on 29 May 2026 and says it runs at more than 100 tokens per second. xAI has not published benchmark scores for it yet.
  • GLM 5.2 is the only open-weights model here, released by Z.ai on 16 June 2026 under the MIT license. You can call it through the API or run it on your own servers if your code must not leave your infrastructure. Z.ai reports a score of 81.0 on Terminal-Bench 2.1. In our run it was the slowest model and used the most output tokens.

To add other models to the comparison, look up their exact IDs in the AI/ML API model catalog instead of guessing. Names and versions change, and a wrong ID makes the script print an error for that row.

Step 6: Turn It Into a Shell Helper Function

Once the script works, wrap it in a shell function so you can call it from any directory. Add this to your ~/.bashrc:

askall() {  
( cd ~/llm-compare && source .venv/bin/activate && python3 compare_llms.py "$@" )
}

Reload the file with command:

source ~/.bashrc

Now askall "review this regex: ^\d{3}-\d{4}$" sends the question to all 5 models from anywhere in the terminal.

Frequently Asked Questions (FAQ)

Q: Do I need a separate account for each AI provider?

A: No. The script sends every request to one OpenAI-compatible endpoint, so a single AI/ML API key covers all 5 models in this tutorial. You do not need accounts with Anthropic, DeepSeek, xAI or Z.ai, and you only install one Python package, openai.

Q: Why does one model return an ERROR in the output?

A: The most common cause is a wrong or outdated model ID. Check the exact ID on the model’s page in the AI/ML API catalog and update the MODELS list. Other causes are an invalid API key (HTTP 401), an empty balance, or a rate limit (HTTP 429), which usually clears after a short wait.

Q: How do I calculate the cost of one request?

A: Multiply the input tokens by the model’s input price and the output tokens by its output price, then divide each by 1,000,000. For example, in our run DeepSeek V4 Flash used 42 input and 279 output tokens, which cost (42 × $0.39 + 279 × $1.56) / 1,000,000, or about $0.00045.

Q: Can I use this script with other OpenAI-compatible APIs?

A: Yes. The script reads the endpoint from the OPENAI_BASE_URL environment variable. For example, set it to http://localhost:11434/v1 to query local models in Ollama and put their names in MODELS. Ollama ignores the API key, but the script expects one, so set AIMLAPI_KEY to any value.

Q: Can I run these models locally instead?

A: Only some of them. GLM 5.2 is published as open weights under the MIT license, but at 753B parameters it needs a multi-GPU server. Claude Sonnet 5 and Grok Build 0.1 have no public weights, so you can use them only through an API. For local experiments on a single machine, smaller open models in Ollama are a more practical choice.

Q: What is an API aggregator?

A: It's a service that gives you one endpoint and one billing account for many models from different providers. That makes it convenient for testing and comparing LLMs without managing separate accounts. The trade-off is that you depend on a middleman for pricing, uptime, and billing accuracy.

Q: Is AI/ML API safe to use?

A: AI/ML API is a functioning API aggregation service, but developers should evaluate it for their own requirements before relying on it in production. Start small and test the service with your actual workload.

Q: What is the minimum amount I can add?

A: AI/ML API's current documentation lists $20 as the minimum top-up amount. Check the Subscription Plans page for current pricing and options.

Q: Can I get a refund if I don't use my credits?

A: According to AI/ML API's Terms and Conditions, purchases are final and non-refundable. Their Help Center also states that top-ups are non-refundable. Review the Terms and Conditions before adding funds.

Q: Should I enable auto-top-up?

A: It's safer to keep auto-top-up disabled initially. First verify the service, billing accuracy, and your actual usage. You can review the relevant settings in Managing Your Subscription.

Q: How can I verify that I'm being billed correctly?

A: Run a small number of controlled API requests, record the models and token usage, and compare the expected cost with the amount deducted from your account. Don't assume the displayed cost is correct without testing it yourself.

Q: Should I use AI/ML API directly in production?

A: Don't make it your sole production dependency until you've tested its reliability, latency, error rates, and billing behavior under your specific workload. Consider maintaining a fallback provider for critical applications.

Q: What should I check before adding a large amount of money?

A: Check the current pricing, refund policy, auto-top-up settings, model availability, rate limits, API reliability, and billing behavior. Policies and pricing can change, so verify the current documentation before committing significant funds.

Q: Where can I read the official terms?

A: See AI/ML API's Terms and Conditions, Managing Your Subscription, and Subscription Plans.

Conclusion

With one OpenAI-compatible endpoint, a virtual environment and a short Python script, you can test Claude Sonnet 5, DeepSeek V4 Pro and Flash, Grok Build 0.1 and GLM 5.2 on your own prompts from a Linux terminal.

Adding another model to the comparison takes one line in the MODELS list: no new SDK, account or key. Run it on the tasks you do every day, compare the time, the tokens and the answers, and pick the model based on those results.

You May Also Like

Leave a Comment

* By using this form you agree with the storage and handling of your data by this website.

This site uses Akismet to reduce spam. Learn how your comment data is processed.

This website uses cookies to improve your experience. By using this site, we will assume that you're OK with it. Accept Read More