HyperAIHyperAI

MCP Tools Reference

Every tool exposed by the HyperAI MCP server, grouped by feature, with parameters and behavior notes.

The HyperAI MCP server exposes 58 tools across six feature groups. All tools run under your own account after OAuth sign-in, so results are scoped to the containers, projects, and datasets you can access. List tools return 30 items per page.

You normally don't call these tools yourself — your AI assistant picks them based on what you ask for. This page is a reference for understanding what the assistant can and cannot do.

Confirmation before writes

The assistant does not run write operations silently. Before creating a container (compute_create_job), restarting one (compute_restart_workspace), creating a project (compute_create_project), creating a dataset entry or version (dataset_create, dataset_create_version), changing who can see a project or dataset (compute_update_project_permissions, dataset_update_permissions), or deploying, restarting, scaling, or stopping a serving (serving_create, serving_create_version, serving_restart_version, serving_update_version, serving_stop_version), it restates what it is about to do — including the billing consequence — and waits for your explicit go-ahead, even when you named the resource yourself. Issuing a personal access token (user_create_personal_access_token) and deleting a dataset (dataset_delete), a serving version (serving_delete_version), a serving (serving_delete), a container run (compute_delete_job), or a project (compute_delete_project) require your explicit approval of the concrete values; the assistant will not do any of these on "you decide". In batch or "just get it done" flows, the assistant keeps deletions out of the flow and uses reversible actions such as stopping a container or a serving version instead — unless you explicitly named the resource to delete.

User

Tools for querying your own account — profile, quota, billing, usage, and subscriptions — and for managing personal access tokens. They always operate on the signed-in account and take no username parameter.

ToolDescription
user_get_profileGet your account profile
user_get_quotaGet storage quota, prepaid compute minutes, and account limits
user_list_transactionsList billing transactions
user_list_usagesList resource usage records
user_list_subscriptionsList subscriptions
user_get_spend_analysis_guideGet the spend analysis guide
user_list_personal_access_tokensList personal access tokens
user_create_personal_access_tokenCreate a personal access token
user_revoke_personal_access_tokenRevoke a personal access token

user_get_profile

Get your account profile: username, display name, email, registration date, roles, membership status, balance, and the organizations you belong to. An organization's ID (not its display name) is what the username parameter of the compute and dataset tools accepts — both to list that organization's resources (e.g. compute_list_jobs) and to create and manage them under the organization (e.g. compute_create_job). Useful for verifying the connection. For an organization's capability flags, seat quota, and members, use the org_get and org_list_members tools.

No parameters.

user_get_quota

Get your storage quota (used / total / remaining), prepaid compute minutes per resource type, and account limits such as how many containers can run at once per resource and GPU type, and the number of datasets and projects.

Three things are easy to confuse. Balance is money (see user_get_profile). Quota is prepaid subscription minutes plus storage (this tool). Limitations are caps on concurrency and counts (this tool). Zero remaining minutes for a resource is not by itself a blocker — it only means the prepaid subscription is used up, and pay-as-you-go usage still charges your balance for any resource you can afford. Whether a container can actually start is decided at creation time by the checks compute_create_job runs (affordability, cluster capacity, per-resource limits, runtime, project, billing plan); quota is never one of them — with one exception: when the remaining storage quota is 0, creating a container fails with "无存储额度" and nothing is created. Reads are unaffected, so the assistant can still report your balance and quota. If the total storage quota is 0, only more allowance helps (deleting data frees nothing); if the total is non-zero, deleting datasets, versions, or containers frees space.

No parameters.

user_list_transactions

List billing transactions — recharges, charges, subscription renewals, refunds, and transfers — newest first. Each amount carries its own currency.

Parameters:

  • type — filter by direction: recharge (deposits, gifts, vouchers), spend (charges), refund, transfer (transfers between accounts), or all (default)
  • page — page number

user_list_usages

List resource usage records, newest first. Compute records carry a duration; serving records are billed as (end − start) × replica count; storage records carry the size change that counts against your storage quota.

Parameters:

  • page — page number

user_list_subscriptions

List your subscriptions across all categories — membership, storage expansion, prepaid compute, one-time permanent compute purchases, and organization seats — with plan, price per period, validity, and auto-renewal status. An active subscription with auto-renewal enabled is charged automatically at the end of each period.

Canceling a subscription is only possible in the HyperAI console, not through this tool.

No parameters.

user_get_spend_analysis_guide

Returns a guide for analyzing account spending. Assistants should read it before answering spending questions ("why did my balance drop", "what am I paying for"), so they collect the right data from the other user tools first.

No parameters.

user_list_personal_access_tokens

List your personal access tokens (PATs) — long-lived API credentials — with name, creation and expiry dates, and last-used time. The token plaintext is never included: it is shown exactly once, when the token is created. An empty list does not prove no token was ever created (revoked tokens are hidden), and a listed token is not necessarily usable — expired tokens remain in the list, so check expiresAt.

No parameters.

user_create_personal_access_token

Create a personal access token for your own account. A PAT survives password changes and logouts; it stops working only when revoked or expired.

Issuing a token is an approval-tier operation: a PAT is equivalent to your account password, so the assistant restates the token name and its expiry — the expires_at it will pass, or the 90-day default when omitted — and waits for your explicit approval before calling. It will not issue a token on "you decide", nor as a side step of another task.

Parameters:

  • name — a label for the token, at most 30 characters; identification only, it plays no part in authentication
  • expires_at — optional expiry as an ISO 8601 UTC instant (e.g. 2026-11-15T00:00:00Z); must be in the future. Omit for the default of 90 days from creation.

The token is shown only once

The token plaintext appears only in this tool's response and can never be retrieved again. Store it securely right away. If it leaks, revoke it with user_revoke_personal_access_token.

user_revoke_personal_access_token

Revoke a personal access token. After revoking, the tool re-reads the list to verify the token is gone.

Parameters:

Revocation is immediate and permanent

A revoked token stops working right away — anything still using it is cut off — and revocation cannot be undone. The assistant should list your tokens and get your explicit confirmation before calling it.

Compute

Tools for working with compute containers and projects — the same objects you see in the Gear section of the console. Tools that operate on a specific container or project accept an optional username: pass an organization's ID to work on that organization's resources, and pass the same value on every call about the same container.

ToolDescription
compute_list_jobsList compute containers
compute_list_projectsList projects
compute_list_resourcesList available compute resource tiers
compute_list_plansList billing plans
compute_list_runtimesList container runtime environments
compute_get_jobGet container details
compute_get_job_metricsGet container metrics
compute_get_projectGet project details and history
compute_get_job_readmeRead a container's README
compute_get_job_notebookRead a container's notebook
compute_get_create_job_guideGet the container creation workflow guide
compute_create_jobCreate a workspace container
compute_create_projectCreate an empty project
compute_stop_jobStop a running container
compute_restart_workspaceRestart a stopped workspace
compute_update_jobUpdate a container's description
compute_update_projectUpdate project settings
compute_update_project_permissionsChange who can see a project
compute_update_job_portsManage container port mappings
compute_delete_jobPermanently delete a container run
compute_delete_projectPermanently delete a project

compute_list_jobs

List your workspace and task containers, newest first. Each entry is one run of a project.

Parameters:

  • username — optional; an organization you belong to (your organizations are listed by user_get_profile). Omit to list your own containers.
  • status — running / succeeded / failed / cancelled / all (default all)
  • q — substring match on container names
  • page — page number

compute_list_projects

List projects, optionally narrowed to those carrying certain tags.

Parameters:

  • username — optional; an organization you belong to (your organizations are listed by user_get_profile)
  • q — substring match on project names
  • tags — optional; only projects carrying all of these tags (exact platform catalog names, case-sensitive)
  • page — page number

compute_list_resources

List available compute resource tiers (GPU/CPU) with specs, whether your balance can afford each, and current load.

Parameters:

  • username — optional

compute_list_plans

List billing plans: pay-as-you-go and time-boxed packages with durations and prices.

Parameters:

compute_list_runtimes

List available container runtime environments (images), excluding deprecated ones.

Parameters:

  • username — optional

compute_get_job

Get full details of a container, including its access URL. Secret values are masked.

Parameters:

  • job_id_or_url — a container ID or a console container URL
  • username — optional; the organization that owns the container. Omit for your own.

compute_get_job_metrics

Get summarized system and custom metrics of a container (latest / min / max / average).

Parameters:

  • job_id — container ID

compute_get_project

Get project details plus its container execution history (workspace and task runs).

Parameters:

  • project_id_or_url — a project ID or a console URL
  • status — filter the execution history: running / succeeded / failed / cancelled / all
  • page — page number
  • username — optional; the organization that owns the project. Omit for your own.

compute_get_job_readme

Read a container's README as rendered HTML. A container without a README is a normal result — tutorial content usually lives in the notebook instead.

Parameters:

  • job_id_or_url — a container ID, a project ID, or a console URL

compute_get_job_notebook

Read a container's notebook content. Notebook content is only available once the container has stopped: a running container — typically a freshly cloned tutorial — returns nothing even when the notebook is already on disk. The assistant treats that as "not readable yet", not as an empty container, and points you at the container's Jupyter link instead.

Parameters:

  • job_id_or_url — a container ID, a project ID, or a console URL

compute_get_create_job_guide

Returns the recommended step-by-step workflow for creating a container (billing, data binding, environment variables, ports, creating under an organization, cloning a public tutorial). Assistants should read this before calling compute_create_job.

One rule worth knowing: on a brand-new account with no projects, containers, or datasets, a request like "let me try a notebook" makes the assistant recommend cloning a popular public tutorial (via resources_search_public_projects) before creating a blank container.

No parameters.

compute_create_job

Create a new workspace container. Billing starts as soon as the container starts, so the assistant restates the plan (resource, image, project, billing) and waits for your explicit go-ahead before calling — even when you named the resource yourself.

Parameters:

  • resource — required; a resource name from compute_list_resources
  • runtime — required; a runtime name from compute_list_runtimes
  • project_id / new_project_name — exactly one of the two: create in an existing project, or create a new project
  • description — optional container description
  • plan_id — optional; a time-boxed billing plan from compute_list_plans. Omit for pay-as-you-go.
  • auto_renew — whether a time-boxed plan renews automatically (default true)
  • idle_timeout_minutes — auto-stop after idle time (default 30, 0 disables)
  • env — environment variables, a list of {name, value, secret}; value may be text, a number, or a boolean
  • data_bindings — data to mount, a list of {source, mount_path, writable}
  • ports — custom port mappings, a list of {port, name}
  • username — optional; an organization ID (its ID, not its display name) to create the container under that organization. The container is billed to the organization's balance, and every follow-up call about it must pass the same username. Omit to create under your own account.

A few rules the tool enforces, matching the console's behavior:

  • Data bindings mount at /input0 through /input4 (read-only sources such as datasets and models) and /output (only a previous container output of the form <owner>/jobs/<job-id>/output can be bound there). See Data Binding.
  • Environment variables must not use the reserved OPENBAYES_ prefix, and names containing TOKEN / SECRET / KEY / PASSWORD must be marked as secret. See Environment Variables.
  • Environment variable values are sent exactly as written. A number ({ "name": "NCCL_SHM_DISABLE", "value": 1 }) and a string ("value": "1") both reach the container as the usual text (NCCL_SHM_DISABLE=1); they differ only in the type recorded in openbayes_params.json and returned when the container is read back, so pick the type your code expects. Nested values (objects and arrays) are not accepted. Secret values are sent untouched.
  • Port 8080 is reserved for the container's built-in service and cannot be mapped. See Custom Port Mapping.
  • The chosen resource must be affordable with your current balance and not at full load, and the chosen billing plan must belong to that resource.

compute_create_project

Create an empty project — a container of runs with nothing started, free of charge. Use it to set up a project (name, description, tags) ahead of time; compute_create_job with new_project_name creates one implicitly but cannot tag it and always starts a billable container. Every call creates a new project, so the assistant checks compute_list_projects first and never reuses an existing project by name. Before writing, the tool checks the account's project-count limit and validates every tag against the platform catalog. Creating a project is a confirmation-tier operation: the assistant restates the name, owner, and tags and waits for your go-ahead.

Parameters:

  • name — required project name
  • description — optional description
  • tags — optional tag list; existing platform catalog names, case-sensitive
  • username — optional; an organization ID to create the project under that organization

compute_stop_job

Stop a running container.

Parameters:

  • job_id — exact container ID; URLs are rejected
  • username — optional; the organization that owns the container. Omit for your own.

Stopping a container is destructive

compute_stop_job terminates the running workload. Unsaved state outside the persisted working directory is lost, just as when stopping a container from the console. In batch or cleanup flows, stopping is the reversible default — the assistant prefers it over deleting anything (compute_delete_job and compute_delete_project are approval-tier and remove run data for good).

compute_restart_workspace

Restart a stopped workspace container with its previous configuration. Restarting resumes billing at the previous resource's rate, so the assistant restates the container, its resource, and the billing consequence and waits for your explicit go-ahead — even when you asked for the restart yourself.

Parameters:

  • job_id — container ID
  • username — optional; the organization that owns the workspace — restarting bills the organization's balance. Omit for your own.

compute_update_job

Change a container's description — a note on one run, such as what it was for or what it produced. Works on running and stopped containers alike; nothing else about the run changes. Zero-cost and reversible, so it needs no confirmation. Success is judged by the description reading back as written.

Parameters:

  • job_id — exact container ID; URLs are rejected
  • description — the new description (an empty string clears it)
  • username — optional; the organization that owns the container. Omit for your own.

compute_update_project

Update project settings: name, description, idle timeout, and tags. Who can see the project is a separate tool, compute_update_project_permissions.

Parameters:

  • project_id_or_url — a project ID or a console URL
  • name — optional new name
  • description — optional new description
  • idle_timeout_minutes — optional; applies project-wide, 0 disables
  • add_tags / remove_tags — optional tag lists; tags are validated against the platform catalog
  • username — optional; the organization that owns the project. Omit for your own.

compute_update_project_permissions

Change who can see a project (and its containers' notebooks and outputs). visibility is private (owner only), public (anyone on the platform can browse and view it), or public_download (anyone can also bulk-download — only on an explicit ask, and not available for an organization's project). org_access, for an organization-owned project only, is private (members cannot see it), read_download, or write. Pass only what changes; the owner is resolved from the project itself.

The tool enforces the platform's own rules before writing and explains any refusal: org_access only applies to organization-owned projects; public_download is refused on them; and an organization's project cannot be public while org_access is private — the tool names the parameter to add instead of raising the other level on its own. It also verifies that the levels the platform returns match what was requested. Making a project public does not list it on the public pages — listing is a separate, tag-driven admin flow.

Changing visibility is a confirmation-tier operation: the assistant restates the current and new levels (both levels for an organization-owned project) and waits for your go-ahead, even when you named the level yourself.

Parameters:

  • project_id_or_url — a project ID or a console URL
  • visibility — optional; private / public / public_download
  • org_access — optional, organization-owned projects only; private / read_download / write

compute_update_job_ports

Add or remove custom port mappings on a container. A mapped port gets a publicly reachable URL, so opening a port publishes whatever listens on it to the internet — the assistant should confirm with you before opening one. Port 8080 is reserved and cannot be mapped.

Parameters:

  • job_id — container ID
  • add_ports — a list of {port, name}
  • remove_ports — a list of port numbers to remove
  • username — optional; the organization that owns the container. Omit for your own.

compute_delete_job

Permanently delete one run (container) of a project together with its /output data. The storage it used is freed immediately, and anything that mounted this run's output (other containers, servings) loses that data; the project and its other runs stay. The run must have ended (CANCELLED, SUCCEEDED, or FAILED) — a running one has to be stopped first. After deleting, the tool re-reads the run to verify it is gone.

Parameters:

  • job_id — exact container ID; URLs are rejected
  • confirm_name — must match the run's exact name as compute_list_jobs shows it, e.g. my-project/3
  • username — optional; the organization that owns the container. Omit for your own.

Deletion cannot be undone

The assistant reads the run with compute_get_job, restates the target (ID, name, project, output size), and gets your explicit approval before calling. Deletion is an approval-tier operation: in batch or "just get it done" flows the assistant keeps it out of the flow and stops containers instead — unless you explicitly named the run to delete.

compute_delete_project

Permanently delete a project together with every run in it and their /output data — anything that mounted those outputs loses the data. No run may be running: stop them first with compute_stop_job and wait for CANCELLED. After deleting, the tool re-reads the project to verify it is gone.

Parameters:

  • project_id_or_url — a project ID or a console URL
  • confirm_name — must match the project's exact current name; checked by the tool and again by the platform

Deletion cannot be undone

The assistant reads the project with compute_get_project, restates the target (ID, name, number of runs), and gets your explicit approval before calling. Deletion is an approval-tier operation: in batch or "just get it done" flows the assistant keeps it out of the flow and stops containers instead — unless you explicitly named the project to delete.

Resources

Tools for discovering public content on the platform.

ToolDescription
resources_search_public_projectsSearch public projects

resources_search_public_projects

Search public projects (tutorials and community projects), filterable by tags. Each result carries a clone count as a popularity signal, and a public project can be cloned as the starting point of a new container.

Parameters:

  • q — search keywords
  • tags — a list of tag names to filter by
  • sort — result order: LAST_ACTIVE_AT_DESC (default) / LAST_ACTIVE_AT_ASC / CLONE_COUNT_DESC / CLONE_COUNT_ASC; use CLONE_COUNT_DESC to surface the most popular projects
  • page — page number

Dataset

Tools for managing datasets and models — searching, inspecting, creating entries and versions, updating metadata, changing visibility, and deleting. Uploading data is not covered here: that happens in the web console.

ToolDescription
dataset_searchSearch datasets and models
dataset_getGet dataset or model details
dataset_createCreate an empty dataset or model entry
dataset_updateUpdate dataset or model metadata
dataset_update_permissionsChange who can see a dataset or model
dataset_create_versionAdd an empty version to a dataset or model
dataset_deletePermanently delete a dataset or model

Search datasets and models — your own and public ones — in one call. Each result includes the binding name to use as a data_bindings source in compute_create_job.

Parameters:

  • q — search keywords
  • category — dataset / model / all (default all)
  • username — optional; search under an organization instead of your own account
  • tags — a list of tag names to filter public results by; tags are a platform-controlled vocabulary matched exactly
  • page — page number for your own results
  • public_page — page number for public results

dataset_get

Get the full details of a dataset or model in one call: metadata, owner, permissions, every version, and the selected version's README as rendered HTML. Each version carries the binding name to use as a data_bindings source in compute_create_job; the top-level size spans all versions.

Parameters:

  • dataset_id_or_url — a dataset/model ID or a console URL; a trailing version segment in the URL is honored
  • version — optional version number; overrides the version in the URL

dataset_create

Create a new, empty dataset or model entry — a metadata shell with no versions and no data. Every call creates a new entry, so check with dataset_search first to avoid duplicates. The assistant restates the entry (name, kind, owner, tags) and waits for your explicit go-ahead before calling.

Parameters:

  • name — required entry name
  • kind — dataset / model (default dataset)
  • description — optional description
  • tags — optional tag list; tags must be existing names from the platform tag catalog (case-sensitive)
  • username — optional; create under an organization instead of your own account

dataset_update

Update a dataset's or model's metadata: name, description, kind, and tags. Pass only the fields to change. Organization-owned entries work directly — the owner is resolved automatically. Who can see the entry is a separate tool, dataset_update_permissions.

Parameters:

  • dataset_id_or_url — a dataset/model ID or a console URL
  • name — optional new name
  • description — optional new description
  • kind — optional; change between dataset and model
  • add_tags / remove_tags — optional tag lists; tags are validated against the platform catalog before anything is written

dataset_update_permissions

Change who can see a dataset or model. visibility is private (owner only), public (anyone on the platform can browse and view it), or public_download (anyone can also bulk-download it — only on an explicit ask, and not available for an organization's entry). org_access, for an organization-owned entry only, is private (members cannot see it), read_download, or write. Pass only what changes.

The tool enforces the platform's own rules before writing and explains any refusal: org_access only applies to organization-owned entries; public_download is refused on them; and an organization's entry cannot be public while org_access is private — the tool names the parameter to add instead of raising the other level on its own. It also verifies that the levels the platform returns match what was requested. Making an entry public does not list it on the public pages — listing is a separate, tag-driven admin flow.

Changing visibility is a confirmation-tier operation: the assistant restates the current and new levels (both levels for an organization-owned entry) and waits for your go-ahead, even when you named the level yourself.

Parameters:

  • dataset_id_or_url — a dataset/model ID or a console URL
  • visibility — optional; private / public / public_download
  • org_access — optional, organization-owned entries only; private / read_download / write

dataset_create_version

Add a new, empty version to an existing dataset or model — the next version number, with no data in it yet. Data goes into it afterwards through the console; an empty version cannot be mounted into a container until it has content. The tool refuses when the latest version is itself still empty (fill that one first) and returns the new version's binding name. Creating a version is a confirmation-tier operation: the assistant restates the entry and its current latest version and waits for your go-ahead.

Parameters:

  • dataset_id_or_url — a dataset/model ID or a console URL

dataset_delete

Permanently delete a dataset or model, including every version and all uploaded data. After deleting, the tool re-reads the entry to verify it is gone.

Parameters:

  • dataset_id_or_url — a dataset/model ID or a console URL
  • confirm_name — must match the entry's exact current name

Deletion cannot be undone

dataset_delete removes every version and all uploaded data permanently. The assistant should read the entry with dataset_get and get your explicit approval before calling it. Deletion is an approval-tier operation: in batch or "just get it done" flows the assistant keeps it out of the flow and leaves the candidates to you — unless you explicitly named the entry to delete.

Org

Tools for querying the organizations you belong to. All three are read-only — inviting members, changing roles, and leaving an organization are console operations. An organization's ID doubles as the username parameter of the compute and dataset tools — to list that organization's resources (compute_list_jobs, dataset_search) and to create and manage them under the organization (compute_create_job, dataset_create).

ToolDescription
org_listList your organizations
org_getGet organization details
org_list_membersList organization members

org_list

List the organizations you belong to, with your role in each (OWNER / MEMBER / PENDING) and per-organization capability flags such as canCreateProject and canCreateInvitation. Whether you can do something in an organization is answered by the capability flags, not the role.

No parameters.

org_get

Get one organization's details: profile (display name, description, type, locked state), seat quota (used / total / remaining), your capability flags in that organization, and the organization's resource limits — how many containers can run at once per resource and GPU type, dataset and project counts, and upload size. These limits, not your personal ones from user_get_quota, govern anything created under the organization with username.

Parameters:

  • org_id — the organization's ID, not its display name

org_list_members

List an organization's members, 30 per page: username, display name, role, and join date, plus whether you can remove each member or change their role. Rows with the PENDING role are outstanding invitations.

Parameters:

  • org_id — the organization's ID
  • page — page number

Serving

Tools for deploying models as inference servings and managing them — the same objects as the Model Deployment section of the console. A serving is a stable endpoint; each deployment behind it is an immutable version (runtime, resource, replica count, and data bindings are frozen when the version is created). Only one version of a serving runs at a time: starting or restarting a version stops the running one. Tools that operate on a specific serving accept an optional username: pass an organization's ID to work on that organization's servings, and pass the same value on every call about the same serving.

Serving must be enabled on your account; if it is not, serving_list and serving_create say so instead of returning an empty list or a permission error.

ToolDescription
serving_listList servings
serving_getGet serving details and readiness
serving_get_metricsGet a version's resource metrics
serving_get_telemetryGet a version's request telemetry
serving_get_deploy_guideGet the deployment workflow guide
serving_createCreate a serving and deploy its first version
serving_create_versionDeploy a new version of a serving
serving_restart_versionBring a stopped version back with its frozen configuration
serving_update_versionScale a live version or annotate it
serving_updateUpdate serving settings and access control
serving_stop_versionStop a running version
serving_delete_versionPermanently delete a version
serving_deletePermanently delete a serving
serving_list_api_keysList serving API keys
serving_create_api_keyIssue a serving API key
serving_update_api_keyRename a key or change its scope
serving_revoke_api_keyRevoke a serving API key

serving_list

List your servings, newest first, with each serving's active versions (status, replica count, resource) and latest version. Link names: frontend is the console page, http the service endpoint, and llm marks an OpenAI-compatible endpoint (the same URL as http, not a second one).

Parameters:

  • username — optional; an organization you belong to. Omit to list your own servings.
  • status — running / stopped / failed / all (default all); running also covers servings that are starting or shutting down
  • q — substring match on serving names
  • page — page number

serving_get

Get one serving's details plus one of its versions: status and sub-status, replica count, resource, runtime, instances and their readiness, data bindings, endpoint links, and a control-plane readiness verdict. By default the version shown is the one currently running — which after a rollback may be older than the latest — or the latest when nothing is running; the response says which selection was made and carries the latest version number.

Readiness here is control-plane only: RUNNING alone does not prove the service answers, so the assistant calls the http link to confirm the data plane. needAuthorization being true means callers must send an API key. Sub-status QUOTA_EXHAUSTED means the platform stopped the version over an exhausted balance.

Parameters:

  • serving_id_or_url — a serving ID or a console serving URL
  • version — optional version number to show instead of the default
  • username — optional; the organization that owns the serving. Omit for your own.

serving_get_metrics

Get a version's resource utilization — CPU (user / system), memory, and per-GPU utilization and VRAM — summarized per instance as latest / average / maximum over a time window. Defaults to the last 30 minutes of the default version at 60-second buckets. Values are as the platform reports them (memory and VRAM in bytes, CPU in millicores).

Parameters:

  • serving_id_or_url — a serving ID or a console URL
  • version — optional version number; omit for the latest
  • from / to — the window: a relative offset (-30m, -24h, -3d), a millisecond timestamp, or now (defaults -30m and now)
  • resolution — bucket size such as 60s, 16m, 1h, 1d (default 60s)
  • username — optional; the organization that owns the serving

Time formats

from and to accept only relative offsets, millisecond timestamps, and now. The platform silently treats any other format — an ISO date in particular — as the current time and returns nothing, so the tool rejects such values up front. An empty result for a running serving usually means the window is too narrow: widen it to -24h before concluding there is no data.

serving_get_telemetry

Get a version's request telemetry, one signal per call: requests (request count and duration distribution per bucket), latency / ttft / total_tokens (LLM request latency, time to first token, and tokens per request — for servings whose links include llm), events (per-request event rows), and llm_events (per-LLM-request token and latency rows). Distribution signals return count-weighted summaries plus the raw buckets.

For an LLM serving the llm signals are the primary ones — the general requests and events streams can be empty while the LLM ones are healthy, so empty general telemetry is not proof of no traffic.

Parameters:

  • serving_id_or_url — a serving ID or a console URL
  • signal — required; requests / latency / ttft / total_tokens / events / llm_events
  • version — optional version number; omit for the latest
  • from / to — the window, same formats as serving_get_metrics (defaults -30m and now)
  • resolution — bucket size for the distribution signals (default 60s); ignored for events
  • tail — for events / llm_events: how many of the most recent rows to return (default 100)
  • method — for requests: the HTTP method to filter on (default POST)
  • username — optional; the organization that owns the serving

serving_get_deploy_guide

Returns the step-by-step guide for deploying a model as a serving: what the working directory must carry (start.sh with the OPENBAYES_SERVING_PRODUCTION port switch), the recommended test-in-a-workspace-first flow, how serving_create derives a version's configuration from a source container, verification, API keys, and the stop → delete order. Assistants read it first whenever you ask to deploy a model, turn a container into an API endpoint, roll out a new version, or are unsure whether you need a restart or a new version.

No parameters.

serving_create

Create a serving and deploy its first version. Billing starts when the version starts (resource × replica count), so the assistant restates the plan — resource, runtime, bindings, replica count, and that billing starts — and waits for your explicit go-ahead, even when you named the container or resource yourself.

There are two ways to describe the version:

  • From a tested workspace container — pass source_job_id. The tool locks the resource to the container's, matches its runtime in the catalog, keeps the container's input bindings read-only, and mounts the container's own output at /output. data_bindings may then only add extra /input* mounts; an /output entry is refused.
  • Explicitly — pass resource, runtime, and data_bindings yourself, with exactly one binding at /output.

Either way, whatever is mounted at /output must carry a start.sh that switches its port on the OPENBAYES_SERVING_PRODUCTION environment variable (8080 in a workspace, 80 in a serving) and listens on 0.0.0.0. The tool cannot look inside a container or a dataset, so the assistant confirms this with you before calling. See Getting Started with Model Deployment for the script.

Parameters:

  • name — required serving name
  • description — optional description
  • source_job_id — optional; the container to deploy from
  • resource — a resource name from compute_list_resources; omitted with source_job_id (inherited)
  • runtime — a runtime name from compute_list_runtimes; with source_job_id, omit to inherit or pass to migrate
  • data_bindings — read-only mounts, a list of {data, mount_path} where data is a dataset's binding name from dataset_search or a container output <owner>/jobs/<job-id>/output, and mount_path is /input0 … /input4 or /output
  • replica — number of instances (default 1); billing scales with it
  • username — optional; an organization ID to create the serving under that organization, billed to its balance. Every follow-up call must pass the same value.

Before writing, the tool checks that serving is enabled, the resource exists, is affordable, and is not at full load, the account's per-resource parallel limit has a free slot (a serving takes the same slot as a container — if the limit is full, stop the source workspace first), and the runtime is available.

serving_create_version

Deploy a new version of an existing serving — for a changed workspace (new code or model in /output), a different runtime or resource, or a different binding set. Billing starts when the version starts, and the serving's currently running version is stopped. Configuration comes from source_job_id or from explicit resource + runtime + data_bindings, with the same rules and checks as serving_create. The assistant restates the plan, including that the running version stops, and waits for your explicit go-ahead.

This is not the tool for bringing a stopped version back with the same configuration — that is serving_restart_version. Minting a new version for a configuration-identical restart wastes a version number and breaks consumers pinned to the old one.

Parameters:

  • serving_id_or_url — a serving ID or a console URL
  • source_job_id / resource / runtime / data_bindings — as in serving_create
  • replica — number of instances; defaults to the latest version's
  • description — optional version note, for example what changed
  • username — optional; the organization that owns the serving

serving_restart_version

Bring a stopped (CANCELLED) or FAILED version back online with its frozen configuration — same runtime, resource, bindings, and replica count, and no new version number. This is what "restart the serving", "it was shut down, bring it back", or "start the cancelled version again" means. Restarting resumes billing, and if another version of the serving is running it is stopped, so the assistant restates the serving, version, resource, replica count, and consequences, and waits for your explicit go-ahead.

Only a CANCELLED or FAILED version can be restarted: a RUNNING one needs nothing, a PREPARING one is already starting, and a CANCELLING one must reach CANCELLED first — the tool says which applies.

Parameters:

  • serving_id_or_url — a serving ID or a console URL
  • version — optional version number; omit for the latest
  • username — optional; the organization that owns the serving — restarting bills the organization's balance

serving_update_version

Change the only mutable settings of a version: replica (scale a live version up or down — billing scales with it) and description (annotate the version). Runtime, resource, and bindings are frozen per version — changing those is serving_create_version. A replica change on a stopped version is refused with a pointer to serving_restart_version. A replica change is a confirmation-tier operation; a description-only change needs no confirmation.

Parameters:

  • serving_id_or_url — a serving ID or a console URL
  • version — optional version number; omit for the latest
  • replica — new number of instances for a live version
  • description — new version note
  • username — optional; the organization that owns the serving

serving_update

Change a serving's own settings: name, description, and need_authorization — the access-control switch. Turning it on makes every call without a valid API key fail (issue keys with serving_create_api_key); turning it off makes the endpoint publicly callable. Changes take up to 5 minutes to apply. See API Key Management. Changing need_authorization is a confirmation-tier operation; a name or description change needs no confirmation.

Parameters:

  • serving_id_or_url — a serving ID or a console URL
  • name — optional new name
  • description — optional new description
  • need_authorization — optional; true requires an API key, false makes the endpoint public
  • username — optional; the organization that owns the serving

serving_stop_version

Stop a running version. The endpoint for that version goes down and callers fail immediately; billing for it stops. The version keeps its configuration and can be brought back with serving_restart_version (a cold start; in-memory state is lost). Without version, the tool stops the version that is actually running.

Parameters:

  • serving_id_or_url — a serving ID or a console URL
  • version — optional version number; omit for the running one
  • username — optional; the organization that owns the serving

Stopping is the reversible default

In batch or cleanup flows, the assistant stops versions rather than deleting anything, and leaves deletion candidates to you.

serving_delete_version

Permanently delete one version of a serving. Only a CANCELLED or FAILED version can be deleted — stop a running one first. After deleting, the tool re-reads the version to verify it is gone.

Parameters:

  • serving_id_or_url — a serving ID or a console URL
  • version — required; the exact version number to delete
  • username — optional; the organization that owns the serving

Deletion cannot be undone

The assistant reads the version with serving_get, restates the target (serving name, version number, status), and gets your explicit approval before calling. Deletion is an approval-tier operation: in batch or "just get it done" flows the assistant keeps it out of the flow — unless you explicitly named the version to delete.

serving_delete

Permanently delete a serving — its endpoint, configuration, and every version. No version may be running: stop them first and wait for CANCELLED. After deleting, the tool re-reads the serving to verify it is gone.

Parameters:

  • serving_id_or_url — a serving ID or a console URL
  • confirm_name — must match the serving's exact current name; checked by the tool and again by the platform
  • username — optional; the organization that owns the serving

Deletion cannot be undone

The assistant reads the serving with serving_get, restates the target (ID, name, version count), and gets your explicit approval before calling. Deletion is an approval-tier operation: in batch or "just get it done" flows the assistant keeps it out of the flow — unless you explicitly named the serving to delete.

serving_list_api_keys

List the API keys issued for calling servings: record ID, name, scope, bound servings, creation and last-use times. The secret itself is never available after creation — the list carries no key value, and a lost key means issuing a new one and revoking the old. The record ID is what serving_update_api_key and serving_revoke_api_key take.

The scope is inferred from the bound servings, because the platform does not report a key's type: a key with bound servings is serving-scoped, and a key with none shows as account — which is also what a serving-scoped key looks like after its servings were deleted.

Parameters:

  • page — page number
  • username — optional; the organization whose keys to list

serving_create_api_key

Issue an API key for calling servings whose access control is on. Scope serving binds the key to the listed servings; scope account opens every serving of the account. A new key takes up to 5 minutes to become effective. The assistant restates the key name and scope and waits for your explicit go-ahead.

Parameters:

  • name — required key name
  • scope — required; serving or account
  • serving_ids — the servings the key opens; required for scope serving
  • username — optional; an organization ID for a key on that organization's servings

The key is shown only once

The key appears only in this tool's response and can never be retrieved again — the assistant shows it in full and asks you to save it. Because the platform returns only the secret, the tool identifies the new record by comparing the key list before and after creation; if it cannot, it reports the record ID as unknown rather than guessing, and it warns when another key already has the same name.

serving_update_api_key

Rename a key or change its scope — which servings it opens, or whether it opens every serving of the account. The secret never changes. Narrowing the scope cuts off callers on the removed servings within up to 5 minutes. A scope change is a confirmation-tier operation; a rename needs no confirmation.

A key with bound servings can be renamed with scope omitted. A key with none — account-wide, or serving-scoped with its servings since deleted — needs scope stated explicitly: the tool never assumes account on its own, because that would silently open every serving.

Parameters:

  • id — the key record's ID from serving_list_api_keys, not the secret
  • name — optional new name
  • scope — optional; serving or account
  • serving_ids — the full new set of servings for scope serving
  • username — optional; the organization that owns the key

serving_revoke_api_key

Revoke a serving API key. Every caller still using it is cut off within up to 5 minutes. After revoking, the tool re-reads the list to verify the key is gone.

Parameters:

  • id — the key record's ID from serving_list_api_keys, not the secret
  • confirm_name — must match the key's exact name
  • username — optional; the organization that owns the key

Revocation cannot be undone

The assistant lists your keys, restates the target by ID, name, and the servings it opens, and gets your explicit confirmation before calling.