llm-api.arc.vt.edu

Description

https://llm-api.arc.vt.edu/api/v1/ provides OpenAI and Anthropic compatible API endpoints to a selection of LLMs hosted and run by ARC. It is based on the Open WebUI platform, and integrates several inference and integration capabilities, including retrieval-augmented generation (RAG), web search, vision, image generation, and text embeddings.

Access

  • All Virginia Tech students, faculty, and staff may access the service at https://llm-api.arc.vt.edu/api/v1/ using a personal API key. No separate ARC account is required.

  • There is no charge to individual users for accessing the hosted models via the API.

  • Users must generate an API key through https://llm.arc.vt.edu User profile > Settings > Account > API keys. Keys are unique to each user and must be kept confidential.

Caution

Sharing API keys with other users is strictly prohibited. Any unauthorized use of shared credentials will be reported to the IT Security Office and access will be terminated.

Restrictions

Data classification: Researchers can use this tool for high-risk data. This service is approved by VT Information Technology Security Office (ITSO) for processing sensitive or regulated data. However, researchers are reminded to consult with VT Privacy and Research Data Protection Program (PRDP) and the Office of Export and Secure Research Compliance regarding the storage and analysis of high-risk data to comply with specific regulations. Note that some high-risk data (e.g. data regulated by DFARS, ITAR, etc.) require additional protections and the LLM might not be approved for use with those data types.

Usage limits

To ensure fair access for all users, the ARC LLM service enforces the following limits on both the API and web interface:

Limit

Value

Description

Maximum concurrent requests

Per model

Concurrency varies per model; see the model table. On models with a fixed number, additional requests are rejected until an active request completes. On fair-queued models, additional requests wait in a shared queue instead (see Fair-queued models).
Solution: For fixed-number models, limit your application to no more than the stated concurrent requests.

Maximum requests in flight per user

20

Across all models combined, counting both running and queued requests. Additional requests are rejected with HTTP 429 until one completes.
Solution: Send no more than 20 requests at once in total, even when you use several models.

Maximum tokens per non-streaming request

8,000

Non-streaming requests with long outputs are more likely to time out.
Solution: Enable streaming for requests that generate longer outputs.

Fair-queued models

GLM-5.3 doesn’t give each user a fixed number of request slots. Every request enters a shared queue, and the service decides what runs next based on current load:

  • When the model is lightly used, your requests start right away, and one user can use most of the model’s capacity.

  • When the model is busy, capacity is split fairly among the people using it right now. Users who have recently used more (long prompts, long outputs) wait longer than users who have used less. One large job can’t lock anyone else out.

  • Chat in the web interface goes first. When the model is saturated, requests from llm.arc.vt.edu are served ahead of API requests. API requests still make progress, but they wait longer at busy times.

  • Your own requests start in the order you sent them.

What this means for your application:

Limit

Value

Description

Requests in flight per user

20

The overall limit above still applies. Up to that limit, requests that the model can’t run right away wait on the server instead of being rejected, so at busy times the first token may take longer to arrive.
Solution: You don’t need to hold yourself to a small concurrency, but stay within 20 in total.

Maximum time per request

30 min

Includes time spent waiting in the queue.
Solution: Set your client timeout long enough to cover queue time, and use streaming for long outputs.

Maximum context

128k tokens

Input plus output. Streaming requests have no separate output limit.

A fair-queued model returns an error in these cases, with the reason in error.code in the response body:

  • HTTP 429 (user_concurrent, queue_full): you already have 20 requests in flight, or the queue is full. Wait a few seconds and retry.

  • HTTP 413 (buffered_max_tokens_exceeded, replica_capacity_exceeded): the request is too large to run. Either you asked for more than 8,000 output tokens without streaming, or the input plus requested output doesn’t fit. Retrying won’t help. Enable streaming, or shorten the input or max_tokens.

Embeddings

The embeddings endpoint (/api/v1/embeddings) is available through the API only and has separate per-user limits:

Limit

Value

Description

Embedding tokens per user

150,000/min

Up to 300,000 tokens at once, refilled at 150,000 tokens per minute.
Solution: Send smaller batches or spread the job over time.

Concurrent embedding requests per user

4

Additional requests are rejected until an active request completes.
Solution: Limit your application to 4 embedding requests at once.

A rejected embedding request returns HTTP 429 with the reason (error.reason) and the number of seconds to wait (error.retry_after_s) in the response body. The response does not include a Retry-After header. See Embedding a large collection for an example that handles this.

These limits are intended to support fair sharing of the service. If your workload requires higher throughput or different models, consider running models on ARC HPC resources using vLLM, Llama.cpp, or the Open OnDemand LLM app.

Models

ARC currently runs several state-of-the-art models. ARC will add or remove models and scale instances dynamically to respond to user demand. You may select your preferred model in the request settings, e.g. "model": "gpt-oss-120b".

  • OpenAI gpt-oss-120b (see model card on Hugging Face). OpenAI’s flagship open-weight model, optimized for fast, high-quality responses across a wide range of general-purpose tasks.

  • Z.AI GLM-5.3 (see model card on Hugging Face). State-of-the-art reasoning model that excels at complex problem solving, coding, mathematics, and scientific tasks.

  • DeepSeek DeepSeek-V4.1-Flash (see model card on Hugging Face). High-performance model optimized for speed and long-context processing, making it well suited for large documents, codebases, and retrieval-augmented applications. Multimodal architecture.

Models also offer pre-configured versions with different levels of thinking modes. You may use:

Model

Parameters

Max context

Concurrency

gpt-oss-120b

reasoning_effort: medium

128k

10

gpt-oss-120b-thinking-low

reasoning_effort: low

128k

gpt-oss-120b-thinking-high

reasoning_effort: high

128k

GLM-5.3

reasoning_effort: max

128k

Fair-queued

GLM-5.3-thinking-high

reasoning_effort: high

128k

DeepSeek-V4.1-Flash

reasoning_effort: high

1M

10

DeepSeek-V4.1-Flash-thinking-low

reasoning_effort: low

1M

DeepSeek-V4.1-Flash-thinking-max

reasoning_effort: max

1M

Embedding models

  • Qwen Qwen3-Embedding-4B (see model card on Hugging Face). Converts text into 2560-dimensional vectors for semantic search, clustering, and retrieval-augmented applications. Each input can be up to 8,192 tokens. Use it with the /api/v1/embeddings endpoint; it does not answer chat requests and is not available in the web interface.

For retrieval, prefix each query (not each document) with a one-line task description. Qwen reports that leaving it off lowers retrieval quality by about 1–5%. See the semantic search example.

Security

This service is hosted entirely on-premises within the ARC infrastructure. No data is sent to any third party outside of the university. All user interactions are logged and preserved in compliance with VT IT Security Office Data Protection policies.

Disclaimer

ARC has implemented safeguards to mitigate the risk of generating unlawful, harmful, or otherwise inappropriate content. Despite these measures, LLMs may still produce inaccurate, misleading, biased, or harmful information. Use of this service is undertaken entirely at the user’s own discretion and risk. The service is provided “as is”, and, to the fullest extent permitted by applicable law, ARC and VT expressly disclaim all warranties, whether express or implied, as well as any liability for damages, losses, or adverse consequences that may result from the use of, or reliance upon, the outputs generated by the models. By using this service, the user acknowledges and accepts these conditions, and agrees to comply with all applicable terms and conditions governing the use of the hosted models, associated software, and underlying platforms.

Examples

Please read the OpenAI API documentation for a comprehensive guide to understand the different ways to interact with the LLM. You may also consult the Open WebUI documentation for API endpoints for additional examples involving Retrieval Augmented Generation (RAG), knowledge collections, image generation, tool calling, web search, etc.

Shell API

Use this API to interact with the LLMs directly from the command line.

OpenAI chat completions

Submit a query to a model.

curl -X POST "https://llm-api.arc.vt.edu/api/v1/chat/completions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-oss-120b",
        "messages": [{
           "role":"user",
           "content":"Why is the sky blue?"
        }]
      }'

Anthropic messages

Submit a query to a model.

curl -X POST "https://llm-api.arc.vt.edu/api/v1/messages" \
  -H "x-api-key: $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-oss-120b",
        "max_tokens": 1024,
        "messages": [{
           "role":"user",
           "content":"Why is the sky blue?"
        }]
      }'

Embeddings

Convert text into vectors. Pass a list in input to embed several texts in one request.

API_KEY="sk-YOUR-API-KEY"

curl -X POST "https://llm-api.arc.vt.edu/api/v1/embeddings" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "Qwen3-Embedding-4B",
        "input": [
           "The capital of France is Paris.",
           "Photosynthesis converts sunlight into chemical energy."
        ]
      }'

Document upload

Upload a document to the LLM. Every file is assigned a unique file id. You can use the file ids to do Retrieval Augmented Generation (RAG).

API_KEY="sk-YOUR-API-KEY"

curl -X POST \
  -H "Authorization: Bearer $API_KEY" \
  -H "Accept: application/json" \
  -F "file=@/path/to/file.pdf" https://llm-api.arc.vt.edu/api/v1/files/

Retrieval Augmented Generation (RAG)

Upload a file, extract its file id, and submit a query about the document to the LLM.

API_KEY="sk-YOUR-API-KEY"

## Upload document and get file ID
file_id=$(curl -s -X POST \
  -H "Authorization: Bearer $API_KEY" \
  -H "Accept: application/json" \
  -F "file=@document.pdf" \
  https://llm-api.arc.vt.edu/api/v1/files/ | jq -r '.id')

## Use the file ID in the request
request=$(jq -n \
  --arg model "gpt-oss-120b" \
  --arg file_id "$file_id" \
  --arg prompt "Create a summary of the document" \
  '{
    model: $model,
    messages: [{role: "user", content: $prompt}],
    files: [{type: "file", id: $file_id}]
  }')

## Make the chat completion request with the file
curl -X POST "https://llm-api.arc.vt.edu/api/v1/chat/completions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d "$request"

Reasoning effort

You may change the reasoning effort on many models, e.g. gpt-oss-120b to (low, medium (default), high).

API_KEY="sk-YOUR-API-KEY"

curl -X POST "https://llm-api.arc.vt.edu/api/v1/chat/completions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "gpt-oss-120b",
        "messages": [{
           "role":"user",
           "content":"Why is the sky blue?"
        }],
        "reasoning_effort": "high"
      }'

Image generation

Caution

Image generation and editing use Qwen-Image-2.1, released under the Qwen Research License, which permits use of the model for non-commercial research and evaluation. Qwen has clarified that generated images are not covered by the license and that you retain the rights to the content you generate. Virginia Tech’s Standard for Acceptable Use separately prohibits using university systems for commercial purposes. If you use generated images to create, train, or improve an AI model that you distribute or make available, the license requires you to display “Built with Qwen” or “Improved using Qwen” in that model’s documentation.

This approach generates an image using Qwen/Qwen-Image-2.1.

API_KEY="sk-YOUR-API-KEY"
OUTPUT="output.png"

RESPONSE=$(curl -s -X POST "https://llm-api.arc.vt.edu/api/v1/images/generations" \
   -H "Authorization: Bearer $API_KEY" \
   -H "Content-Type: application/json" \
   -d '{       
       "prompt": "Generate a picture of a white lab dog",
       "size": "512x512"
      }'
)

# Extract returned URL
FILE_URL=$(echo "$RESPONSE" | jq -r '.[0].url')

# Download image
curl -s -L -o "$OUTPUT" \
  -H "Authorization: Bearer $API_KEY" \
  "https://llm-api.arc.vt.edu$FILE_URL"

Image edition

This approach edits an image using Qwen/Qwen-Image-2.1, the same model used for generation. See the license restrictions.

API_KEY="sk-YOUR-API-KEY"
INPUT="input.png"
OUTPUT="output.png"

# Encode input image
IMG_B64=$(base64 -w0 "$INPUT")

# Submit the image edit request
RESPONSE=$(curl -s -X POST "https://llm-api.arc.vt.edu/api/v1/images/edit" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d @- <<EOF
{
  "image": "data:image/png;base64,$IMG_B64",
  "prompt": "Change the color of the red blanket to blue"
}
EOF
)

# Extract returned URL
FILE_URL=$(echo "$RESPONSE" | jq -r '.[0].url')

# Download image
curl -s -L -o "$OUTPUT" \
  -H "Authorization: Bearer $API_KEY" \
  "https://llm-api.arc.vt.edu$FILE_URL"

Python API

Certain libraries are required for the API use in Python such as openai and requests. You may install them using pip:

pip install openai requests

Chat completions

Submit a query to a model.

from openai import OpenAI
import argparse

# Modify OpenAI's API key and API base to use the server.
openai_api_key = "sk-YOUR-API-KEY"
openai_api_base = "https://llm-api.arc.vt.edu/api/v1"

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "What is Virginia Tech known for?"},
]

client = OpenAI(
    api_key=openai_api_key,
    base_url=openai_api_base,
)

chat_completion = client.chat.completions.create(
    model="gpt-oss-120b",
    messages=messages,
)

print(chat_completion)

Embeddings

Convert a batch of texts into vectors.

from openai import OpenAI

openai_api_key = "sk-YOUR-API-KEY"
openai_api_base = "https://llm-api.arc.vt.edu/api/v1"

client = OpenAI(
    api_key=openai_api_key,
    base_url=openai_api_base,
)

response = client.embeddings.create(
    model="Qwen3-Embedding-4B",
    input=[
        "The capital of France is Paris.",
        "Photosynthesis converts sunlight into chemical energy.",
    ],
)

for item in response.data:
    print(item.index, len(item.embedding), item.embedding[:3])
print(response.usage)

Embedding a large collection

Embed many texts in batches, and wait when the service returns a rate-limit error. The server puts the wait time in the error body, so this example turns off the client’s built-in retries and uses that value instead.

import time
from openai import OpenAI, RateLimitError

openai_api_key = "sk-YOUR-API-KEY"
openai_api_base = "https://llm-api.arc.vt.edu/api/v1"
BATCH_SIZE = 32

client = OpenAI(
    api_key=openai_api_key,
    base_url=openai_api_base,
    max_retries=0,
)

def embed_batch(texts):
    while True:
        try:
            response = client.embeddings.create(model="Qwen3-Embedding-4B", input=texts)
            return [item.embedding for item in response.data]
        except RateLimitError as e:
            error = e.body if isinstance(e.body, dict) else {}
            wait = error.get("retry_after_s", 5)
            print(f"Rate limited ({error.get('reason')}), waiting {wait}s")
            time.sleep(wait)

texts = [f"Document number {i}" for i in range(1000)]  # replace with your texts

embeddings = []
for start in range(0, len(texts), BATCH_SIZE):
    embeddings.extend(embed_batch(texts[start:start + BATCH_SIZE]))

print(f"Embedded {len(embeddings)} texts")

Document upload

Upload a document to the LLM. Every file is assigned a unique file id. You can use the file ids to do Retrieval Augmented Generation (RAG).

import os
import requests

api_key="sk-YOUR-API-KEY"
base_url="https://llm-api.arc.vt.edu/api/v1/files/"
file_path="document.pdf"

if os.path.isfile(file_path):
    with open(file_path, "rb") as file:
        response = requests.post(
            base_url,
            headers={
                "Authorization": f"Bearer {api_key}",
                "Accept": "application/json",
            },
            files={"file": file},
        )        
        if response.status_code == 200:
            print(f"Uploaded {file_path} successfully!")
        else:
            print(f"Failed to upload {file_path}. Status code: {response.status_code}")
else:
    print(f"File not found")

Retrieval Augmented Generation (RAG)

Upload a file, extract its file id, and submit a query about the document to the LLM.

import os
import requests
import json

api_key="sk-YOUR-API-KEY"
file_path="document.pdf"

def upload_file(file_path):
    if not os.path.isfile(file_path):
        raise FileNotFoundError(f"File not found: {file_path}")

    with open(file_path, "rb") as file:
        response = requests.post(
            "https://llm-api.arc.vt.edu/api/v1/files/",
            headers={
                "Authorization": f"Bearer {api_key}",
                "Accept": "application/json",
            },
            files={"file": file},
        )

    if response.status_code == 200:
        data = response.json()
        file_id = data.get("id")
        if file_id:
            print(f"Uploaded {file_path} successfully! File ID: {file_id}")
            return file_id
        else:
            raise RuntimeError("Upload succeeded but no file id returned.")
    else:
        raise RuntimeError(f"Failed to upload {file_path}. Status code: {response.status_code}")

file_id = upload_file(file_path)

url = "https://llm-api.arc.vt.edu/api/v1/chat/completions"
headers = {
    "Authorization": f"Bearer {api_key}",
    "Content-Type": "application/json"
    }
data = {
    "model": "gpt-oss-120b",
    "messages": [{
        "role": "user",
        "content": "Create a summary of the document"}],
    "files": [{"type": "file", "id": file_id}],
}

response = requests.post(url, headers=headers, data=json.dumps(data))
print(response.text)

Image generation

This approach generates an image using Qwen/Qwen-Image-2.1. See the license restrictions.

import base64
import requests
from urllib.parse import urlparse
from openai import OpenAI

openai_api_key = "sk-YOUR-API-KEY"
openai_api_base = "https://llm-api.arc.vt.edu/api/v1"

client = OpenAI(
    api_key=openai_api_key,
    base_url=openai_api_base,
)

response = client.images.generate(
    prompt="A gray tabby cat hugging an otter with an orange scarf",
    size="512x512",
)

base_url = urlparse(openai_api_base)
image_url = f"{base_url.scheme}://{base_url.netloc}" + response[0].url
headers = {"Authorization": f"Bearer {openai_api_key}"}
img_data = requests.get(image_url, headers=headers).content

with open("output.png", 'wb') as handler:
    handler.write(img_data)

Image edition

This approach edits an image using Qwen/Qwen-Image-2.1, the same model used for generation. See the license restrictions.

import base64
import json
import requests
from pathlib import Path

BASE_URL = "https://llm-api.arc.vt.edu/api/v1"
API_KEY = "sk-YOUR-API-KEY"
EDIT_INSTRUCTION = "Change the color of the orange scarf to blue."
INPUT_IMAGE = "input.png"
OUTPUT_IMAGE = "output.png"

def convert_image_to_base64(image_path: str) -> str:
    image_path = Path(image_path)
    if not image_path.exists():
        raise FileNotFoundError(f"Image file not found: {image_path}")
    
    return base64.b64encode(image_path.read_bytes()).decode("utf-8")

def request_image_edit(edit_instruction: str, image_path: str) -> list:
    print("Submitting request for image edit...")

    url = f"{BASE_URL}/images/edit"
    headers = {
        "Authorization": f"Bearer {API_KEY}",
        "Content-Type": "application/json",
    }

    image_b64 = convert_image_to_base64(image_path)

    payload = {
        "form_data": { 
            "prompt": edit_instruction,
            "image": f"data:image/png;base64,{image_b64}"
        },
    }

    resp = requests.post(url, headers=headers, json=payload, timeout=300)
    resp.raise_for_status()
    return resp.json()   # <-- returns a LIST

def extract_file_url(result_list: list) -> str:
    """
    The API returns:  [ { "url": "/api/v1/files/.../content" } ]
    """
    if not isinstance(result_list, list) or not result_list:
        raise ValueError("Unexpected response: expected non-empty list")

    entry = result_list[0]

    if "url" not in entry:
        raise ValueError(f"No URL found in response: {entry}")

    return entry["url"]

def download_file_from_url(file_url: str, out_path: str) -> None:
    headers = {"Authorization": f"Bearer {API_KEY}"}

    # If URL is relative, prepend host
    if file_url.startswith("/"):
        file_url = BASE_URL + file_url.replace("/api/v1", "")

    r = requests.get(file_url, headers=headers, timeout=60)
    r.raise_for_status()

    Path(out_path).write_bytes(r.content)
    print(f"Saved edited image to {out_path}")

result = request_image_edit(EDIT_INSTRUCTION, INPUT_IMAGE)
file_url = extract_file_url(result)
download_file_from_url(file_url, OUTPUT_IMAGE)

Image to text Python API

import requests
import base64
from pathlib import Path

url = "https://llm-api.arc.vt.edu/api/v1/chat/completions"
openai_api_key = "sk-YOUR-API-KEY"
image_path = "bonnie.jpg"

def convert_image_to_base64(image_path: str) -> str:
    image_path = Path(image_path)
    if not image_path.exists():
        raise FileNotFoundError(f"Image file not found: {image_path}")
    
    with open(image_path, "rb") as img_file:
        encoded = base64.b64encode(img_file.read()).decode("utf-8")
    return encoded

headers = {
    "Authorization": f"Bearer {openai_api_key}",
    "Content-Type": "application/json"
}

image_b64 = convert_image_to_base64(image_path)

data = {
    "model": "DeepSeek-V4.1-Flash",
    "messages": [
        {
            "role": "user",
            "content": [
                {
                    "type": "text",
                    "text": "Describe the image"
                },
                {
                    "type": "image_url",
                    "image_url": {"url": f"data:image/jpeg;base64,{image_b64}"}
                }
            ]
        }
    ]
}

response = requests.post(url, headers=headers, json=data, timeout=30)
print(response.json()["choices"][0]["message"]["content"])

For repeated queries on the same image, see the reusable uploaded image workflow below.

Video analysis

Important

The LLM API might not currently accept native video_url input. For visual video analysis, use this video decomposition example, it samples representative frames from a video and sends them to the LLM API.

This approach analyzes the visual content only; audio from the video is not processed. If your workflow requires combined video and audio understanding, consider getting a GPU on the cluster.

Vision workflow with reusable uploaded images

For repeated vision queries on the same image, This workflow is supported where the image is uploaded once through /api/v1/files/ and then referenced by file ID in subsequent requests.

This differs from the inline base64 workflow, where the full image must be included in every request.

When to use this workflow

Use this approach when:

  • You need to run multiple prompts against the same image

  • You want to avoid repeatedly sending the same image in the request payload

For one-off vision requests, the inline base64 workflow might still be simpler.

Endpoint note: /api/v1 vs /api

ARC exposes two API styles that are used differently in this workflow:

  • /api/v1/... is the documented OpenAI-compatible path shown in the ARC examples

  • /api/... is the native Open WebUI path. In this example, /api/chat/completions is used for the file ID based vision request.

Follow the endpoint shown in the sample codes for the specific workflow you are using.

Example: upload an image once and reuse it by file ID
"""Upload an image once and reuse it in a vision request via VT ARC LLM API."""

import os
import requests

API_KEY = "sk-YOUR-API-KEY"
BASE = "https://llm-api.arc.vt.edu"

# 1. Upload the image 
with open("input.png", "rb") as f:
    upload_resp = requests.post(
        f"{BASE}/api/v1/files/",
        headers={"Authorization": f"Bearer {API_KEY}"},
        files={"file": ("input.png", f)},
        timeout=120,
    )

upload_resp.raise_for_status()
file_id = upload_resp.json()["id"]

print(f"File ID: {file_id}")

# 2. Send chat completion referencing the file ID
#    NOTE: use /api/chat/completions (not /api/v1/) — the native Open WebUI
#    endpoint resolves bare file IDs to images server-side.
chat_resp = requests.post(
    f"{BASE}/api/chat/completions",
    headers={
        "Authorization": f"Bearer {API_KEY}",
        "Content-Type": "application/json",
    },
    json={
        "model": "DeepSeek-V4.1-Flash",
        "messages": [
            {
                "role": "user", "content": [
                    {"type": "text", "text": "Describe this image in detail."},
                    {"type": "image_url", "image_url": {"url": file_id}},
                ],
            }
        ],
    },
    timeout=120,
)

chat_resp.raise_for_status()
resp = chat_resp.json()

print(resp["choices"][0]["message"]["content"])