MCP Tools Reference
Every tool exposed by the HyperAI MCP server, grouped by feature, with parameters and behavior notes.
The HyperAI MCP server exposes 58 tools across six feature groups. All tools run under your own account after OAuth sign-in, so results are scoped to the containers, projects, and datasets you can access. List tools return 30 items per page.
You normally don't call these tools yourself — your AI assistant picks them based on what you ask for. This page is a reference for understanding what the assistant can and cannot do.
Confirmation before writes
The assistant does not run write operations silently. Before creating a container (compute_create_job), restarting one (compute_restart_workspace), creating a project (compute_create_project), creating a dataset entry or version (dataset_create, dataset_create_version), changing who can see a project or dataset (compute_update_project_permissions, dataset_update_permissions), or deploying, restarting, scaling, or stopping a serving (serving_create, serving_create_version, serving_restart_version, serving_update_version, serving_stop_version), it restates what it is about to do — including the billing consequence — and waits for your explicit go-ahead, even when you named the resource yourself. Issuing a personal access token (user_create_personal_access_token) and deleting a dataset (dataset_delete), a serving version (serving_delete_version), a serving (serving_delete), a container run (compute_delete_job), or a project (compute_delete_project) require your explicit approval of the concrete values; the assistant will not do any of these on "you decide". In batch or "just get it done" flows, the assistant keeps deletions out of the flow and uses reversible actions such as stopping a container or a serving version instead — unless you explicitly named the resource to delete.
User
Tools for querying your own account — profile, quota, billing, usage, and subscriptions — and for managing personal access tokens. They always operate on the signed-in account and take no username parameter.
| Tool | Description |
|---|---|
user_get_profile | Get your account profile |
user_get_quota | Get storage quota, prepaid compute minutes, and account limits |
user_list_transactions | List billing transactions |
user_list_usages | List resource usage records |
user_list_subscriptions | List subscriptions |
user_get_spend_analysis_guide | Get the spend analysis guide |
user_list_personal_access_tokens | List personal access tokens |
user_create_personal_access_token | Create a personal access token |
user_revoke_personal_access_token | Revoke a personal access token |
user_get_profile
Get your account profile: username, display name, email, registration date, roles, membership status, balance, and the organizations you belong to. An organization's ID (not its display name) is what the username parameter of the compute and dataset tools accepts — both to list that organization's resources (e.g. compute_list_jobs) and to create and manage them under the organization (e.g. compute_create_job). Useful for verifying the connection. For an organization's capability flags, seat quota, and members, use the org_get and org_list_members tools.
No parameters.
user_get_quota
Get your storage quota (used / total / remaining), prepaid compute minutes per resource type, and account limits such as how many containers can run at once per resource and GPU type, and the number of datasets and projects.
Three things are easy to confuse. Balance is money (see user_get_profile). Quota is prepaid subscription minutes plus storage (this tool). Limitations are caps on concurrency and counts (this tool). Zero remaining minutes for a resource is not by itself a blocker — it only means the prepaid subscription is used up, and pay-as-you-go usage still charges your balance for any resource you can afford. Whether a container can actually start is decided at creation time by the checks compute_create_job runs (affordability, cluster capacity, per-resource limits, runtime, project, billing plan); quota is never one of them — with one exception: when the remaining storage quota is 0, creating a container fails with "无存储额度" and nothing is created. Reads are unaffected, so the assistant can still report your balance and quota. If the total storage quota is 0, only more allowance helps (deleting data frees nothing); if the total is non-zero, deleting datasets, versions, or containers frees space.
No parameters.
user_list_transactions
List billing transactions — recharges, charges, subscription renewals, refunds, and transfers — newest first. Each amount carries its own currency.
Parameters:
type— filter by direction:recharge(deposits, gifts, vouchers),spend(charges),refund,transfer(transfers between accounts), orall(default)page— page number
user_list_usages
List resource usage records, newest first. Compute records carry a duration; serving records are billed as (end − start) × replica count; storage records carry the size change that counts against your storage quota.
Parameters:
page— page number
user_list_subscriptions
List your subscriptions across all categories — membership, storage expansion, prepaid compute, one-time permanent compute purchases, and organization seats — with plan, price per period, validity, and auto-renewal status. An active subscription with auto-renewal enabled is charged automatically at the end of each period.
Canceling a subscription is only possible in the HyperAI console, not through this tool.
No parameters.
user_get_spend_analysis_guide
Returns a guide for analyzing account spending. Assistants should read it before answering spending questions ("why did my balance drop", "what am I paying for"), so they collect the right data from the other user tools first.
No parameters.
user_list_personal_access_tokens
List your personal access tokens (PATs) — long-lived API credentials — with name, creation and expiry dates, and last-used time. The token plaintext is never included: it is shown exactly once, when the token is created. An empty list does not prove no token was ever created (revoked tokens are hidden), and a listed token is not necessarily usable — expired tokens remain in the list, so check expiresAt.
No parameters.
user_create_personal_access_token
Create a personal access token for your own account. A PAT survives password changes and logouts; it stops working only when revoked or expired.
Issuing a token is an approval-tier operation: a PAT is equivalent to your account password, so the assistant restates the token name and its expiry — the expires_at it will pass, or the 90-day default when omitted — and waits for your explicit approval before calling. It will not issue a token on "you decide", nor as a side step of another task.
Parameters:
name— a label for the token, at most 30 characters; identification only, it plays no part in authenticationexpires_at— optional expiry as an ISO 8601 UTC instant (e.g.2026-11-15T00:00:00Z); must be in the future. Omit for the default of 90 days from creation.
The token is shown only once
The token plaintext appears only in this tool's response and can never be retrieved again. Store it securely right away. If it leaks, revoke it with user_revoke_personal_access_token.
user_revoke_personal_access_token
Revoke a personal access token. After revoking, the tool re-reads the list to verify the token is gone.
Parameters:
id— the token's ID fromuser_list_personal_access_tokens, not the token plaintextconfirm_name— must match the token's exact name
Revocation is immediate and permanent
A revoked token stops working right away — anything still using it is cut off — and revocation cannot be undone. The assistant should list your tokens and get your explicit confirmation before calling it.
Compute
Tools for working with compute containers and projects — the same objects you see in the Gear section of the console. Tools that operate on a specific container or project accept an optional username: pass an organization's ID to work on that organization's resources, and pass the same value on every call about the same container.
| Tool | Description |
|---|---|
compute_list_jobs | List compute containers |
compute_list_projects | List projects |
compute_list_resources | List available compute resource tiers |
compute_list_plans | List billing plans |
compute_list_runtimes | List container runtime environments |
compute_get_job | Get container details |
compute_get_job_metrics | Get container metrics |
compute_get_project | Get project details and history |
compute_get_job_readme | Read a container's README |
compute_get_job_notebook | Read a container's notebook |
compute_get_create_job_guide | Get the container creation workflow guide |
compute_create_job | Create a workspace container |
compute_create_project | Create an empty project |
compute_stop_job | Stop a running container |
compute_restart_workspace | Restart a stopped workspace |
compute_update_job | Update a container's description |
compute_update_project | Update project settings |
compute_update_project_permissions | Change who can see a project |
compute_update_job_ports | Manage container port mappings |
compute_delete_job | Permanently delete a container run |
compute_delete_project | Permanently delete a project |
compute_list_jobs
List your workspace and task containers, newest first. Each entry is one run of a project.
Parameters:
username— optional; an organization you belong to (your organizations are listed byuser_get_profile). Omit to list your own containers.status—running/succeeded/failed/cancelled/all(defaultall)q— substring match on container namespage— page number
compute_list_projects
List projects, optionally narrowed to those carrying certain tags.
Parameters:
username— optional; an organization you belong to (your organizations are listed byuser_get_profile)q— substring match on project namestags— optional; only projects carrying all of these tags (exact platform catalog names, case-sensitive)page— page number
compute_list_resources
List available compute resource tiers (GPU/CPU) with specs, whether your balance can afford each, and current load.
Parameters:
username— optional
compute_list_plans
List billing plans: pay-as-you-go and time-boxed packages with durations and prices.
Parameters:
resource— optional; a resource name fromcompute_list_resources
compute_list_runtimes
List available container runtime environments (images), excluding deprecated ones.
Parameters:
username— optional
compute_get_job
Get full details of a container, including its access URL. Secret values are masked.
Parameters:
job_id_or_url— a container ID or a console container URLusername— optional; the organization that owns the container. Omit for your own.
compute_get_job_metrics
Get summarized system and custom metrics of a container (latest / min / max / average).
Parameters:
job_id— container ID
compute_get_project
Get project details plus its container execution history (workspace and task runs).
Parameters:
project_id_or_url— a project ID or a console URLstatus— filter the execution history:running/succeeded/failed/cancelled/allpage— page numberusername— optional; the organization that owns the project. Omit for your own.
compute_get_job_readme
Read a container's README as rendered HTML. A container without a README is a normal result — tutorial content usually lives in the notebook instead.
Parameters:
job_id_or_url— a container ID, a project ID, or a console URL
compute_get_job_notebook
Read a container's notebook content. Notebook content is only available once the container has stopped: a running container — typically a freshly cloned tutorial — returns nothing even when the notebook is already on disk. The assistant treats that as "not readable yet", not as an empty container, and points you at the container's Jupyter link instead.
Parameters:
job_id_or_url— a container ID, a project ID, or a console URL
compute_get_create_job_guide
Returns the recommended step-by-step workflow for creating a container (billing, data binding, environment variables, ports, creating under an organization, cloning a public tutorial). Assistants should read this before calling compute_create_job.
One rule worth knowing: on a brand-new account with no projects, containers, or datasets, a request like "let me try a notebook" makes the assistant recommend cloning a popular public tutorial (via resources_search_public_projects) before creating a blank container.
No parameters.
compute_create_job
Create a new workspace container. Billing starts as soon as the container starts, so the assistant restates the plan (resource, image, project, billing) and waits for your explicit go-ahead before calling — even when you named the resource yourself.
Parameters:
resource— required; a resource name fromcompute_list_resourcesruntime— required; a runtime name fromcompute_list_runtimesproject_id/new_project_name— exactly one of the two: create in an existing project, or create a new projectdescription— optional container descriptionplan_id— optional; a time-boxed billing plan fromcompute_list_plans. Omit for pay-as-you-go.auto_renew— whether a time-boxed plan renews automatically (defaulttrue)idle_timeout_minutes— auto-stop after idle time (default30,0disables)env— environment variables, a list of{name, value, secret};valuemay be text, a number, or a booleandata_bindings— data to mount, a list of{source, mount_path, writable}ports— custom port mappings, a list of{port, name}username— optional; an organization ID (its ID, not its display name) to create the container under that organization. The container is billed to the organization's balance, and every follow-up call about it must pass the sameusername. Omit to create under your own account.
A few rules the tool enforces, matching the console's behavior:
- Data bindings mount at
/input0through/input4(read-only sources such as datasets and models) and/output(only a previous container output of the form<owner>/jobs/<job-id>/outputcan be bound there). See Data Binding. - Environment variables must not use the reserved
OPENBAYES_prefix, and names containingTOKEN/SECRET/KEY/PASSWORDmust be marked as secret. See Environment Variables. - Environment variable values are sent exactly as written. A number (
{ "name": "NCCL_SHM_DISABLE", "value": 1 }) and a string ("value": "1") both reach the container as the usual text (NCCL_SHM_DISABLE=1); they differ only in the type recorded inopenbayes_params.jsonand returned when the container is read back, so pick the type your code expects. Nested values (objects and arrays) are not accepted. Secret values are sent untouched. - Port 8080 is reserved for the container's built-in service and cannot be mapped. See Custom Port Mapping.
- The chosen resource must be affordable with your current balance and not at full load, and the chosen billing plan must belong to that resource.
compute_create_project
Create an empty project — a container of runs with nothing started, free of charge. Use it to set up a project (name, description, tags) ahead of time; compute_create_job with new_project_name creates one implicitly but cannot tag it and always starts a billable container. Every call creates a new project, so the assistant checks compute_list_projects first and never reuses an existing project by name. Before writing, the tool checks the account's project-count limit and validates every tag against the platform catalog. Creating a project is a confirmation-tier operation: the assistant restates the name, owner, and tags and waits for your go-ahead.
Parameters:
name— required project namedescription— optional descriptiontags— optional tag list; existing platform catalog names, case-sensitiveusername— optional; an organization ID to create the project under that organization
compute_stop_job
Stop a running container.
Parameters:
job_id— exact container ID; URLs are rejectedusername— optional; the organization that owns the container. Omit for your own.
Stopping a container is destructive
compute_stop_job terminates the running workload. Unsaved state outside the persisted working directory is lost, just as when stopping a container from the console. In batch or cleanup flows, stopping is the reversible default — the assistant prefers it over deleting anything (compute_delete_job and compute_delete_project are approval-tier and remove run data for good).
compute_restart_workspace
Restart a stopped workspace container with its previous configuration. Restarting resumes billing at the previous resource's rate, so the assistant restates the container, its resource, and the billing consequence and waits for your explicit go-ahead — even when you asked for the restart yourself.
Parameters:
job_id— container IDusername— optional; the organization that owns the workspace — restarting bills the organization's balance. Omit for your own.
compute_update_job
Change a container's description — a note on one run, such as what it was for or what it produced. Works on running and stopped containers alike; nothing else about the run changes. Zero-cost and reversible, so it needs no confirmation. Success is judged by the description reading back as written.
Parameters:
job_id— exact container ID; URLs are rejecteddescription— the new description (an empty string clears it)username— optional; the organization that owns the container. Omit for your own.
compute_update_project
Update project settings: name, description, idle timeout, and tags. Who can see the project is a separate tool, compute_update_project_permissions.
Parameters:
project_id_or_url— a project ID or a console URLname— optional new namedescription— optional new descriptionidle_timeout_minutes— optional; applies project-wide,0disablesadd_tags/remove_tags— optional tag lists; tags are validated against the platform catalogusername— optional; the organization that owns the project. Omit for your own.
compute_update_project_permissions
Change who can see a project (and its containers' notebooks and outputs). visibility is private (owner only), public (anyone on the platform can browse and view it), or public_download (anyone can also bulk-download — only on an explicit ask, and not available for an organization's project). org_access, for an organization-owned project only, is private (members cannot see it), read_download, or write. Pass only what changes; the owner is resolved from the project itself.
The tool enforces the platform's own rules before writing and explains any refusal: org_access only applies to organization-owned projects; public_download is refused on them; and an organization's project cannot be public while org_access is private — the tool names the parameter to add instead of raising the other level on its own. It also verifies that the levels the platform returns match what was requested. Making a project public does not list it on the public pages — listing is a separate, tag-driven admin flow.
Changing visibility is a confirmation-tier operation: the assistant restates the current and new levels (both levels for an organization-owned project) and waits for your go-ahead, even when you named the level yourself.
Parameters:
project_id_or_url— a project ID or a console URLvisibility— optional;private/public/public_downloadorg_access— optional, organization-owned projects only;private/read_download/write
compute_update_job_ports
Add or remove custom port mappings on a container. A mapped port gets a publicly reachable URL, so opening a port publishes whatever listens on it to the internet — the assistant should confirm with you before opening one. Port 8080 is reserved and cannot be mapped.
Parameters:
job_id— container IDadd_ports— a list of{port, name}remove_ports— a list of port numbers to removeusername— optional; the organization that owns the container. Omit for your own.
compute_delete_job
Permanently delete one run (container) of a project together with its /output data. The storage it used is freed immediately, and anything that mounted this run's output (other containers, servings) loses that data; the project and its other runs stay. The run must have ended (CANCELLED, SUCCEEDED, or FAILED) — a running one has to be stopped first. After deleting, the tool re-reads the run to verify it is gone.
Parameters:
job_id— exact container ID; URLs are rejectedconfirm_name— must match the run's exact name ascompute_list_jobsshows it, e.g.my-project/3username— optional; the organization that owns the container. Omit for your own.
Deletion cannot be undone
The assistant reads the run with compute_get_job, restates the target (ID, name, project, output size), and gets your explicit approval before calling. Deletion is an approval-tier operation: in batch or "just get it done" flows the assistant keeps it out of the flow and stops containers instead — unless you explicitly named the run to delete.
compute_delete_project
Permanently delete a project together with every run in it and their /output data — anything that mounted those outputs loses the data. No run may be running: stop them first with compute_stop_job and wait for CANCELLED. After deleting, the tool re-reads the project to verify it is gone.
Parameters:
project_id_or_url— a project ID or a console URLconfirm_name— must match the project's exact current name; checked by the tool and again by the platform
Deletion cannot be undone
The assistant reads the project with compute_get_project, restates the target (ID, name, number of runs), and gets your explicit approval before calling. Deletion is an approval-tier operation: in batch or "just get it done" flows the assistant keeps it out of the flow and stops containers instead — unless you explicitly named the project to delete.
Resources
Tools for discovering public content on the platform.
| Tool | Description |
|---|---|
resources_search_public_projects | Search public projects |
resources_search_public_projects
Search public projects (tutorials and community projects), filterable by tags. Each result carries a clone count as a popularity signal, and a public project can be cloned as the starting point of a new container.
Parameters:
q— search keywordstags— a list of tag names to filter bysort— result order:LAST_ACTIVE_AT_DESC(default) /LAST_ACTIVE_AT_ASC/CLONE_COUNT_DESC/CLONE_COUNT_ASC; useCLONE_COUNT_DESCto surface the most popular projectspage— page number
Dataset
Tools for managing datasets and models — searching, inspecting, creating entries and versions, updating metadata, changing visibility, and deleting. Uploading data is not covered here: that happens in the web console.
| Tool | Description |
|---|---|
dataset_search | Search datasets and models |
dataset_get | Get dataset or model details |
dataset_create | Create an empty dataset or model entry |
dataset_update | Update dataset or model metadata |
dataset_update_permissions | Change who can see a dataset or model |
dataset_create_version | Add an empty version to a dataset or model |
dataset_delete | Permanently delete a dataset or model |
dataset_search
Search datasets and models — your own and public ones — in one call. Each result includes the binding name to use as a data_bindings source in compute_create_job.
Parameters:
q— search keywordscategory—dataset/model/all(defaultall)username— optional; search under an organization instead of your own accounttags— a list of tag names to filter public results by; tags are a platform-controlled vocabulary matched exactlypage— page number for your own resultspublic_page— page number for public results
dataset_get
Get the full details of a dataset or model in one call: metadata, owner, permissions, every version, and the selected version's README as rendered HTML. Each version carries the binding name to use as a data_bindings source in compute_create_job; the top-level size spans all versions.
Parameters:
dataset_id_or_url— a dataset/model ID or a console URL; a trailing version segment in the URL is honoredversion— optional version number; overrides the version in the URL
dataset_create
Create a new, empty dataset or model entry — a metadata shell with no versions and no data. Every call creates a new entry, so check with dataset_search first to avoid duplicates. The assistant restates the entry (name, kind, owner, tags) and waits for your explicit go-ahead before calling.
Parameters:
name— required entry namekind—dataset/model(defaultdataset)description— optional descriptiontags— optional tag list; tags must be existing names from the platform tag catalog (case-sensitive)username— optional; create under an organization instead of your own account
dataset_update
Update a dataset's or model's metadata: name, description, kind, and tags. Pass only the fields to change. Organization-owned entries work directly — the owner is resolved automatically. Who can see the entry is a separate tool, dataset_update_permissions.
Parameters:
dataset_id_or_url— a dataset/model ID or a console URLname— optional new namedescription— optional new descriptionkind— optional; change betweendatasetandmodeladd_tags/remove_tags— optional tag lists; tags are validated against the platform catalog before anything is written
dataset_update_permissions
Change who can see a dataset or model. visibility is private (owner only), public (anyone on the platform can browse and view it), or public_download (anyone can also bulk-download it — only on an explicit ask, and not available for an organization's entry). org_access, for an organization-owned entry only, is private (members cannot see it), read_download, or write. Pass only what changes.
The tool enforces the platform's own rules before writing and explains any refusal: org_access only applies to organization-owned entries; public_download is refused on them; and an organization's entry cannot be public while org_access is private — the tool names the parameter to add instead of raising the other level on its own. It also verifies that the levels the platform returns match what was requested. Making an entry public does not list it on the public pages — listing is a separate, tag-driven admin flow.
Changing visibility is a confirmation-tier operation: the assistant restates the current and new levels (both levels for an organization-owned entry) and waits for your go-ahead, even when you named the level yourself.
Parameters:
dataset_id_or_url— a dataset/model ID or a console URLvisibility— optional;private/public/public_downloadorg_access— optional, organization-owned entries only;private/read_download/write
dataset_create_version
Add a new, empty version to an existing dataset or model — the next version number, with no data in it yet. Data goes into it afterwards through the console; an empty version cannot be mounted into a container until it has content. The tool refuses when the latest version is itself still empty (fill that one first) and returns the new version's binding name. Creating a version is a confirmation-tier operation: the assistant restates the entry and its current latest version and waits for your go-ahead.
Parameters:
dataset_id_or_url— a dataset/model ID or a console URL
dataset_delete
Permanently delete a dataset or model, including every version and all uploaded data. After deleting, the tool re-reads the entry to verify it is gone.
Parameters:
dataset_id_or_url— a dataset/model ID or a console URLconfirm_name— must match the entry's exact current name
Deletion cannot be undone
dataset_delete removes every version and all uploaded data permanently. The assistant should read the entry with dataset_get and get your explicit approval before calling it. Deletion is an approval-tier operation: in batch or "just get it done" flows the assistant keeps it out of the flow and leaves the candidates to you — unless you explicitly named the entry to delete.
Org
Tools for querying the organizations you belong to. All three are read-only — inviting members, changing roles, and leaving an organization are console operations. An organization's ID doubles as the username parameter of the compute and dataset tools — to list that organization's resources (compute_list_jobs, dataset_search) and to create and manage them under the organization (compute_create_job, dataset_create).
| Tool | Description |
|---|---|
org_list | List your organizations |
org_get | Get organization details |
org_list_members | List organization members |
org_list
List the organizations you belong to, with your role in each (OWNER / MEMBER / PENDING) and per-organization capability flags such as canCreateProject and canCreateInvitation. Whether you can do something in an organization is answered by the capability flags, not the role.
No parameters.
org_get
Get one organization's details: profile (display name, description, type, locked state), seat quota (used / total / remaining), your capability flags in that organization, and the organization's resource limits — how many containers can run at once per resource and GPU type, dataset and project counts, and upload size. These limits, not your personal ones from user_get_quota, govern anything created under the organization with username.
Parameters:
org_id— the organization's ID, not its display name
org_list_members
List an organization's members, 30 per page: username, display name, role, and join date, plus whether you can remove each member or change their role. Rows with the PENDING role are outstanding invitations.
Parameters:
org_id— the organization's IDpage— page number
Serving
Tools for deploying models as inference servings and managing them — the same objects as the Model Deployment section of the console. A serving is a stable endpoint; each deployment behind it is an immutable version (runtime, resource, replica count, and data bindings are frozen when the version is created). Only one version of a serving runs at a time: starting or restarting a version stops the running one. Tools that operate on a specific serving accept an optional username: pass an organization's ID to work on that organization's servings, and pass the same value on every call about the same serving.
Serving must be enabled on your account; if it is not, serving_list and serving_create say so instead of returning an empty list or a permission error.
| Tool | Description |
|---|---|
serving_list | List servings |
serving_get | Get serving details and readiness |
serving_get_metrics | Get a version's resource metrics |
serving_get_telemetry | Get a version's request telemetry |
serving_get_deploy_guide | Get the deployment workflow guide |
serving_create | Create a serving and deploy its first version |
serving_create_version | Deploy a new version of a serving |
serving_restart_version | Bring a stopped version back with its frozen configuration |
serving_update_version | Scale a live version or annotate it |
serving_update | Update serving settings and access control |
serving_stop_version | Stop a running version |
serving_delete_version | Permanently delete a version |
serving_delete | Permanently delete a serving |
serving_list_api_keys | List serving API keys |
serving_create_api_key | Issue a serving API key |
serving_update_api_key | Rename a key or change its scope |
serving_revoke_api_key | Revoke a serving API key |
serving_list
List your servings, newest first, with each serving's active versions (status, replica count, resource) and latest version. Link names: frontend is the console page, http the service endpoint, and llm marks an OpenAI-compatible endpoint (the same URL as http, not a second one).
Parameters:
username— optional; an organization you belong to. Omit to list your own servings.status—running/stopped/failed/all(defaultall);runningalso covers servings that are starting or shutting downq— substring match on serving namespage— page number
serving_get
Get one serving's details plus one of its versions: status and sub-status, replica count, resource, runtime, instances and their readiness, data bindings, endpoint links, and a control-plane readiness verdict. By default the version shown is the one currently running — which after a rollback may be older than the latest — or the latest when nothing is running; the response says which selection was made and carries the latest version number.
Readiness here is control-plane only: RUNNING alone does not prove the service answers, so the assistant calls the http link to confirm the data plane. needAuthorization being true means callers must send an API key. Sub-status QUOTA_EXHAUSTED means the platform stopped the version over an exhausted balance.
Parameters:
serving_id_or_url— a serving ID or a console serving URLversion— optional version number to show instead of the defaultusername— optional; the organization that owns the serving. Omit for your own.
serving_get_metrics
Get a version's resource utilization — CPU (user / system), memory, and per-GPU utilization and VRAM — summarized per instance as latest / average / maximum over a time window. Defaults to the last 30 minutes of the default version at 60-second buckets. Values are as the platform reports them (memory and VRAM in bytes, CPU in millicores).
Parameters:
serving_id_or_url— a serving ID or a console URLversion— optional version number; omit for the latestfrom/to— the window: a relative offset (-30m,-24h,-3d), a millisecond timestamp, ornow(defaults-30mandnow)resolution— bucket size such as60s,16m,1h,1d(default60s)username— optional; the organization that owns the serving
Time formats
from and to accept only relative offsets, millisecond timestamps, and now. The platform silently treats any other format — an ISO date in particular — as the current time and returns nothing, so the tool rejects such values up front. An empty result for a running serving usually means the window is too narrow: widen it to -24h before concluding there is no data.
serving_get_telemetry
Get a version's request telemetry, one signal per call: requests (request count and duration distribution per bucket), latency / ttft / total_tokens (LLM request latency, time to first token, and tokens per request — for servings whose links include llm), events (per-request event rows), and llm_events (per-LLM-request token and latency rows). Distribution signals return count-weighted summaries plus the raw buckets.
For an LLM serving the llm signals are the primary ones — the general requests and events streams can be empty while the LLM ones are healthy, so empty general telemetry is not proof of no traffic.
Parameters:
serving_id_or_url— a serving ID or a console URLsignal— required;requests/latency/ttft/total_tokens/events/llm_eventsversion— optional version number; omit for the latestfrom/to— the window, same formats asserving_get_metrics(defaults-30mandnow)resolution— bucket size for the distribution signals (default60s); ignored for eventstail— forevents/llm_events: how many of the most recent rows to return (default100)method— forrequests: the HTTP method to filter on (defaultPOST)username— optional; the organization that owns the serving
serving_get_deploy_guide
Returns the step-by-step guide for deploying a model as a serving: what the working directory must carry (start.sh with the OPENBAYES_SERVING_PRODUCTION port switch), the recommended test-in-a-workspace-first flow, how serving_create derives a version's configuration from a source container, verification, API keys, and the stop → delete order. Assistants read it first whenever you ask to deploy a model, turn a container into an API endpoint, roll out a new version, or are unsure whether you need a restart or a new version.
No parameters.
serving_create
Create a serving and deploy its first version. Billing starts when the version starts (resource × replica count), so the assistant restates the plan — resource, runtime, bindings, replica count, and that billing starts — and waits for your explicit go-ahead, even when you named the container or resource yourself.
There are two ways to describe the version:
- From a tested workspace container — pass
source_job_id. The tool locks the resource to the container's, matches its runtime in the catalog, keeps the container's input bindings read-only, and mounts the container's own output at/output.data_bindingsmay then only add extra/input*mounts; an/outputentry is refused. - Explicitly — pass
resource,runtime, anddata_bindingsyourself, with exactly one binding at/output.
Either way, whatever is mounted at /output must carry a start.sh that switches its port on the OPENBAYES_SERVING_PRODUCTION environment variable (8080 in a workspace, 80 in a serving) and listens on 0.0.0.0. The tool cannot look inside a container or a dataset, so the assistant confirms this with you before calling. See Getting Started with Model Deployment for the script.
Parameters:
name— required serving namedescription— optional descriptionsource_job_id— optional; the container to deploy fromresource— a resource name fromcompute_list_resources; omitted withsource_job_id(inherited)runtime— a runtime name fromcompute_list_runtimes; withsource_job_id, omit to inherit or pass to migratedata_bindings— read-only mounts, a list of{data, mount_path}wheredatais a dataset's binding name fromdataset_searchor a container output<owner>/jobs/<job-id>/output, andmount_pathis/input0…/input4or/outputreplica— number of instances (default1); billing scales with itusername— optional; an organization ID to create the serving under that organization, billed to its balance. Every follow-up call must pass the same value.
Before writing, the tool checks that serving is enabled, the resource exists, is affordable, and is not at full load, the account's per-resource parallel limit has a free slot (a serving takes the same slot as a container — if the limit is full, stop the source workspace first), and the runtime is available.
serving_create_version
Deploy a new version of an existing serving — for a changed workspace (new code or model in /output), a different runtime or resource, or a different binding set. Billing starts when the version starts, and the serving's currently running version is stopped. Configuration comes from source_job_id or from explicit resource + runtime + data_bindings, with the same rules and checks as serving_create. The assistant restates the plan, including that the running version stops, and waits for your explicit go-ahead.
This is not the tool for bringing a stopped version back with the same configuration — that is serving_restart_version. Minting a new version for a configuration-identical restart wastes a version number and breaks consumers pinned to the old one.
Parameters:
serving_id_or_url— a serving ID or a console URLsource_job_id/resource/runtime/data_bindings— as inserving_createreplica— number of instances; defaults to the latest version'sdescription— optional version note, for example what changedusername— optional; the organization that owns the serving
serving_restart_version
Bring a stopped (CANCELLED) or FAILED version back online with its frozen configuration — same runtime, resource, bindings, and replica count, and no new version number. This is what "restart the serving", "it was shut down, bring it back", or "start the cancelled version again" means. Restarting resumes billing, and if another version of the serving is running it is stopped, so the assistant restates the serving, version, resource, replica count, and consequences, and waits for your explicit go-ahead.
Only a CANCELLED or FAILED version can be restarted: a RUNNING one needs nothing, a PREPARING one is already starting, and a CANCELLING one must reach CANCELLED first — the tool says which applies.
Parameters:
serving_id_or_url— a serving ID or a console URLversion— optional version number; omit for the latestusername— optional; the organization that owns the serving — restarting bills the organization's balance
serving_update_version
Change the only mutable settings of a version: replica (scale a live version up or down — billing scales with it) and description (annotate the version). Runtime, resource, and bindings are frozen per version — changing those is serving_create_version. A replica change on a stopped version is refused with a pointer to serving_restart_version. A replica change is a confirmation-tier operation; a description-only change needs no confirmation.
Parameters:
serving_id_or_url— a serving ID or a console URLversion— optional version number; omit for the latestreplica— new number of instances for a live versiondescription— new version noteusername— optional; the organization that owns the serving
serving_update
Change a serving's own settings: name, description, and need_authorization — the access-control switch. Turning it on makes every call without a valid API key fail (issue keys with serving_create_api_key); turning it off makes the endpoint publicly callable. Changes take up to 5 minutes to apply. See API Key Management. Changing need_authorization is a confirmation-tier operation; a name or description change needs no confirmation.
Parameters:
serving_id_or_url— a serving ID or a console URLname— optional new namedescription— optional new descriptionneed_authorization— optional;truerequires an API key,falsemakes the endpoint publicusername— optional; the organization that owns the serving
serving_stop_version
Stop a running version. The endpoint for that version goes down and callers fail immediately; billing for it stops. The version keeps its configuration and can be brought back with serving_restart_version (a cold start; in-memory state is lost). Without version, the tool stops the version that is actually running.
Parameters:
serving_id_or_url— a serving ID or a console URLversion— optional version number; omit for the running oneusername— optional; the organization that owns the serving
Stopping is the reversible default
In batch or cleanup flows, the assistant stops versions rather than deleting anything, and leaves deletion candidates to you.
serving_delete_version
Permanently delete one version of a serving. Only a CANCELLED or FAILED version can be deleted — stop a running one first. After deleting, the tool re-reads the version to verify it is gone.
Parameters:
serving_id_or_url— a serving ID or a console URLversion— required; the exact version number to deleteusername— optional; the organization that owns the serving
Deletion cannot be undone
The assistant reads the version with serving_get, restates the target (serving name, version number, status), and gets your explicit approval before calling. Deletion is an approval-tier operation: in batch or "just get it done" flows the assistant keeps it out of the flow — unless you explicitly named the version to delete.
serving_delete
Permanently delete a serving — its endpoint, configuration, and every version. No version may be running: stop them first and wait for CANCELLED. After deleting, the tool re-reads the serving to verify it is gone.
Parameters:
serving_id_or_url— a serving ID or a console URLconfirm_name— must match the serving's exact current name; checked by the tool and again by the platformusername— optional; the organization that owns the serving
Deletion cannot be undone
The assistant reads the serving with serving_get, restates the target (ID, name, version count), and gets your explicit approval before calling. Deletion is an approval-tier operation: in batch or "just get it done" flows the assistant keeps it out of the flow — unless you explicitly named the serving to delete.
serving_list_api_keys
List the API keys issued for calling servings: record ID, name, scope, bound servings, creation and last-use times. The secret itself is never available after creation — the list carries no key value, and a lost key means issuing a new one and revoking the old. The record ID is what serving_update_api_key and serving_revoke_api_key take.
The scope is inferred from the bound servings, because the platform does not report a key's type: a key with bound servings is serving-scoped, and a key with none shows as account — which is also what a serving-scoped key looks like after its servings were deleted.
Parameters:
page— page numberusername— optional; the organization whose keys to list
serving_create_api_key
Issue an API key for calling servings whose access control is on. Scope serving binds the key to the listed servings; scope account opens every serving of the account. A new key takes up to 5 minutes to become effective. The assistant restates the key name and scope and waits for your explicit go-ahead.
Parameters:
name— required key namescope— required;servingoraccountserving_ids— the servings the key opens; required for scopeservingusername— optional; an organization ID for a key on that organization's servings
The key is shown only once
The key appears only in this tool's response and can never be retrieved again — the assistant shows it in full and asks you to save it. Because the platform returns only the secret, the tool identifies the new record by comparing the key list before and after creation; if it cannot, it reports the record ID as unknown rather than guessing, and it warns when another key already has the same name.
serving_update_api_key
Rename a key or change its scope — which servings it opens, or whether it opens every serving of the account. The secret never changes. Narrowing the scope cuts off callers on the removed servings within up to 5 minutes. A scope change is a confirmation-tier operation; a rename needs no confirmation.
A key with bound servings can be renamed with scope omitted. A key with none — account-wide, or serving-scoped with its servings since deleted — needs scope stated explicitly: the tool never assumes account on its own, because that would silently open every serving.
Parameters:
id— the key record's ID fromserving_list_api_keys, not the secretname— optional new namescope— optional;servingoraccountserving_ids— the full new set of servings for scopeservingusername— optional; the organization that owns the key
serving_revoke_api_key
Revoke a serving API key. Every caller still using it is cut off within up to 5 minutes. After revoking, the tool re-reads the list to verify the key is gone.
Parameters:
id— the key record's ID fromserving_list_api_keys, not the secretconfirm_name— must match the key's exact nameusername— optional; the organization that owns the key
Revocation cannot be undone
The assistant lists your keys, restates the target by ID, name, and the servings it opens, and gets your explicit confirmation before calling.