together-hello-world
Build a bounded Together AI chat-completion probe with a current model selected from the live catalog, optional streaming, usage capture, and redacted evidence. Use when proving a new inference connection. Trigger with "Together hello world", "first Together request", or "stream Together chat".
Allowed Tools
Provided by Plugin
together-pack
16 source-grounded operator skills for Together AI inference, batch, fine-tuning, deployment, security, and operations
Installation
This skill is included in the together-pack plugin:
/plugin install together-pack@claude-code-plugins-plus
Click to copy
Instructions
Together AI First Inference
Overview
This skill produces the smallest observable chat-completion probe without freezing a model catalog entry into application policy.
Prerequisites
- A configured
TOGETHER_API_KEYsecret reference - Together Python SDK v2, current TypeScript SDK, or an HTTPS client
- A current chat model resolved from the Together catalog
- A small approved prompt containing no sensitive data
Tool Discipline
Use Read, Glob, and Grep to inspect the client wrapper and configuration. Use WebFetch for the current quickstart, model catalog, and response contract. Use Write or Edit only for the approved sample or test file; never place the key in code.
Current Contract
- Call
client.chat.completions.create()with a Together model ID such as the current quickstart example. - For streaming, iterate chunks and guard for an empty
choicesarray before readingdelta.content. - Capture
finish_reasonandusage; boundmax_tokens, timeout, and retry count. - Re-resolve the model from the live catalog before production use because availability and redirects change.
Authentication
The SDK reads TOGETHER_API_KEY. Direct REST requests send the same project key as a Bearer token to https://api.together.ai/v1/chat/completions. Never return or log the header.
Instructions
- Confirm the requested modality is chat and select a current chat-capable model.
- Create one system message and one sanitized user message with an explicit output bound.
- Choose non-streaming for contract inspection or streaming for latency and chunk handling.
- Validate status, non-empty content, finish reason, model identifier, and usage fields.
- Record latency and request outcome without storing the prompt or response when either is sensitive.
- Remove disposable output and hand off the model-selection rationale.
Approval Boundaries
Do not send customer data, regulated content, or proprietary prompts until data handling and retention are approved. Do not silently switch to a more expensive or materially different model.
Output
Return the client/runtime, resolved model, streaming mode, bounded request parameters, status, finish reason, token usage, latency, and redaction disposition.
Error Handling
| Condition | Response |
|---|---|
401 |
Stop and repair credential injection. |
404 |
Refresh the model catalog and deprecation page. |
429 |
Read dynamic limit headers and retry with bounded jitter. |
503 or 504 |
Retry briefly, then surface overload or timeout evidence. |
Examples
The example below shows the minimum redacted evidence expected from a successful invocation of this operator workflow.
model=catalog-resolved; stream=true; max_tokens=128; status=200; finish=stop; secret=redacted