together-pack
Source-grounded operator workflows for Together AI inference, batch, fine-tuning, v2 Dedicated Model Inference, security, cost, and production operations
Installation
Open Claude Code and run this command:
/plugin install together-pack@claude-code-plugins-plus
Use --global to install for all projects, or --project for current project only.
What It Does
> 16 source-grounded Claude Code skills for current Together AI application and model operations.
Skills (16) plugin-local skills
Gate Together AI client, batch, fine-tuning, and deployment changes with offline contract tests plus a protected bounded live lane.
Analyze and diagnose Together AI authentication, billing, request, model, throttling, overload, batch, fine-tuning, and endpoint failures from redacted evidence.
Prepare, submit, monitor, and disposition a Together AI fine-tuning job using SDK v2, validated training data, explicit cost approval, and separate deployment verification.
Run Together AI asynchronous batch inference from validated JSONL through upload, job polling, output/error download, and custom-id reconciliation.
Reduce Together AI spend using measured token usage, live per-model prices, cached-input evidence, batch discounts, model evaluation, and dedicated break-even analysis.
Deploy and roll back Together AI integrations across serverless inference or v2 Dedicated Model Inference with secret injection, health probes, traffic control, and cost shutdown.
Build a bounded Together AI chat-completion probe with a current model selected from the live catalog, optional streaming, usage capture, and redacted evidence.
Install Together AI SDK v2 and configure a project-scoped API key with least-privilege storage and a read-only model-list verification.
Develop Together AI integrations with a fake transport, recorded response shapes, deterministic assertions, and an opt-in bounded live probe.
Review a Together AI production release across model policy, auth, data controls, dynamic limits, retries, observability, cost, deprecations, asynchronous recovery, and rollback.
Analyze and control Together AI serverless concurrency from dynamic per-model request/token headers, bounded queues, jittered retries, and batch or dedicated alternatives.
Design a production Together AI service with a typed provider boundary, policy-based model routing, serverless and dedicated lanes, batch workers, telemetry, budgets, and reversible degradation.
Encapsulate Together AI SDK v2 behind a typed adapter with catalog-resolved models, bounded retries, streaming normalization, usage capture, and OpenAI-compatible migration seams.
Secure Together AI integrations with project-scoped keys, environment isolation, prompt/output data controls, bounded model behavior, safe logging, rotation, and incident response.
Migrate Together AI Python SDK v1, OpenAI-compatible clients, deprecated models, or legacy dedicated endpoints with inventory, contract tests, canaries, and rollback.
Convert Together AI batch, fine-tuning, upload, and dedicated-deployment job states into idempotent internal events using bounded polling and an optional owned callback.