> For the complete documentation index, see [llms.txt](https://docs.agenticflow.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.agenticflow.ai/changelog/office-hour-54-new-models-safer-actions-and-easier-agent-building.md).

# Office Hour #54: New AI Models, Safer Actions & Easier Agent Building

### Quick Recap

Office Hour #54 expands the model catalog and makes everyday work across AgenticFlow safer, clearer, and easier to complete.

This release includes:

* Six new model choices: Gemini 3.7 Flash, Grok 4.6, Muse Spark 1.2, Qwen 3.8 Max, DeepSeek V4 Pro 0813, and Qwen3.8-27B
* Confirmation before deleting workflows, AI Drive items, MCP servers, Workforces, and connections
* Clearer MCP connection and tool-availability guidance
* More reliable navigation across Library, Datasets, Marketplace, and onboarding
* Easier keyboard and assistive-technology use across agent, workflow, workforce, project, connection, and sharing screens
* Clearer labels and action names when creating, configuring, sharing, and publishing work

🎥 Watch the full session:

{% embed url="<https://www.youtube.com/watch?v=gke9mU10S3w>" %}

***

### What Changed & Why

#### 1. Six New Models Added to the Catalog

The Model Selector now includes six additional choices:

* **Gemini 3.7 Flash**
* **Grok 4.6**
* **Muse Spark 1.2**
* **Qwen 3.8 Max**
* **DeepSeek V4 Pro 0813**
* **Qwen3.8-27B**

These additions give builders more flexibility when choosing a model for agents, workflows, content creation, research, and other day-to-day tasks. You can compare them with your existing models to find the response style, speed, and task fit that works best for each use case.

As with any model change, test important live workflows before switching. Responses can vary by model even when the prompt and tools stay the same.

<figure><img src="https://487764224-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZ3ppnJjAH1qBNXEYnDPA%2Fuploads%2Fgit-blob-cef9cb3347a29a4fcc9abb77882039191fca723f%2Foh54-models-grok.png?alt=media" alt="AgenticFlow model selector showing Grok 4.6"><figcaption><p>Grok 4.6 appears in the live model selector with its provider grouping and input/output limits.</p></figcaption></figure>

<figure><img src="https://487764224-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZ3ppnJjAH1qBNXEYnDPA%2Fuploads%2Fgit-blob-3981827971b73f0efa83b8b1f6616670343bd5f9%2Foh54-models-grok-qwen.png?alt=media" alt="AgenticFlow model selector showing Qwen 3.8 Max and Qwen 3.8 27B"><figcaption><p>Qwen 3.8 Max and Qwen 3.8 27B are listed with their model limits.</p></figcaption></figure>

<figure><img src="https://487764224-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZ3ppnJjAH1qBNXEYnDPA%2Fuploads%2Fgit-blob-77ea60525d522bf30fce4399bbde1344ce13ad39%2Foh54-models-qwen-gemini.png?alt=media" alt="AgenticFlow model selector showing Gemini 3.7 Flash and DeepSeek V4 Pro 0813"><figcaption><p>Gemini 3.7 Flash and DeepSeek V4 Pro 0813 appear in the live catalog.</p></figcaption></figure>

<figure><img src="https://487764224-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZ3ppnJjAH1qBNXEYnDPA%2Fuploads%2Fgit-blob-51e8812febfbdda83385889d1e6e98b0493a07fa%2Foh54-models-deepseek-muse.png?alt=media" alt="AgenticFlow model selector showing DeepSeek V4 Pro 0813 and Muse Spark 1.2"><figcaption><p>DeepSeek V4 Pro 0813 and Muse Spark 1.2 show their published input/output limits.</p></figcaption></figure>

### Model research: choose by job, not by leaderboard

The six additions span three different deployment choices: hosted frontier models, hosted models co-trained with an agent harness, and an open-weight model that can run on a local workstation. The useful question is not “which model is #1?” but “which model gives this workflow the required quality, latency, cost, data boundary, and tool behavior?”

The table below is a practical starting point. It combines the model catalog in AgenticFlow with independent benchmark snapshots and publisher documentation. Re-test any production workflow after changing models because a benchmark score does not guarantee the same result on your prompt, tools, or data.

| Model                    | Best first job                                                                           | Why it is interesting                                                                                                                                                                                                                                                                                | Watch-out                                                                                                                                                        |
| ------------------------ | ---------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Gemini 3.7 Flash**     | High-volume multimodal extraction, classification, and fast workflow steps               | Artificial Analysis reports a high-effort Intelligence Index of 56, about 340 output tokens/second, a 1M-token context, and lower measured cost per task than Gemini 3.6 Flash. See the [AA analysis](https://artificialanalysis.ai/articles/gemini-3-7-time-frontier).                              | Speed and price depend on effort level, cache, and provider route. Measure quality on your own documents before making it the default.                           |
| **Grok 4.6**             | Research, long-running coding, and tasks where strong tool use matters                   | Artificial Analysis reports an Intelligence Index of 61, Terminal-Bench 2.1 at 88.4%, and a 500K context window. It is also #5 in the Arena Code/WebDev snapshot below. See the [AA analysis](https://artificialanalysis.ai/articles/grok-4-6-benchmarks-and-analysis).                              | It is hosted rather than local. Treat live-information behavior, data handling, and cost as part of the design, not as afterthoughts.                            |
| **Muse Spark 1.2**       | Repository work, visual-to-code tasks, and long-horizon coding with persistent subagents | The Muse product is paired with a coding harness designed for long tool sequences. It reaches #2 in Arena’s overall lab snapshot and #22 in the Code/WebDev snapshot below.                                                                                                                          | Model and harness behavior are coupled. A good result in Muse Code may not reproduce with a plain chat completion.                                               |
| **Qwen 3.8 Max**         | Hosted coding, analysis, and multimodal workflows that need a long context               | QwenCloud lists 1M context, tool use, structured output, and a 2T-class mixture-of-experts model; Arena Code ranks it #3 in the dated snapshot below.                                                                                                                                                | Vendor-reported model size and benchmark results are not the same as guaranteed application quality. Check current pricing and limits in the provider console.   |
| **DeepSeek V4 Pro 0813** | Text/code/tool workflows where open weights, long context, or a lower-cost route matters | The DeepSeek API exposes OpenAI-compatible and Anthropic-compatible interfaces, and the family is documented as open-weight/MIT with a 1M context route. Arena Code ranks the 0813 high-effort variant #10 in the snapshot below. See the [DeepSeek updates](https://api-docs.deepseek.com/updates). | Provider routes can expose different revisions, effort levels, and prices. Confirm the exact model ID and evaluate tool-call reliability, not just text quality. |
| **Qwen3.8-27B**          | Private/local coding, knowledge work, and repeatable experiments on a workstation        | A 27B open-weight model is unusually capable for its size. Unsloth reports that a 4-bit build can fit in roughly 17–19GB of memory, putting a useful local experiment within reach of a 24GB-class GPU or Apple-silicon machine. See the [Unsloth guide](https://unsloth.ai/docs/models/qwen3.8).    | A quantized weight-file fit is not the same as a usable 1M context. Leave headroom for KV cache, the operating system, tools, and your runtime.                  |

### Reading the Artificial Analysis view

Artificial Analysis’ Intelligence Index is a composite, not a single test. Version 4.1 weights agentic and knowledge-work evaluations such as GDPval-AA v2, Terminal-Bench 2.1, τ³-Bench Banking, Humanity’s Last Exam, GPQA, and coding/science tasks. The [AA methodology update](https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-1) explains the current mix and why scores can move when the index changes.

The supplied screenshot is useful as a “capability per parameter” conversation starter. It is a dated view of the Artificial Analysis chart, not a live guarantee. In that view, Qwen3.8-27B sits around the low-50s while the much larger Qwen3.8 Max sits near the frontier. That is the important customer insight: a small open model can be unusually attractive when privacy, local control, or predictable cost matters, even when a larger hosted model remains stronger on some tasks.

<figure><img src="https://487764224-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZ3ppnJjAH1qBNXEYnDPA%2Fuploads%2Fgit-blob-66723a7c225031773ee22636e5f83662c2110e59%2Foh54-artificial-analysis-parameters.jpg?alt=media" alt="Artificial Analysis Intelligence Index versus total parameters chart supplied for Office Hour 54"><figcaption><p>Artificial Analysis Intelligence Index versus total parameters, supplied snapshot. Use it to discuss efficiency and the Pareto frontier; verify current values in the <a href="https://artificialanalysis.ai/models">live model explorer</a> before making a procurement decision.</p></figcaption></figure>

#### The Qwen3.8-27B versus Opus 4.6 headline, correctly framed

“Qwen3.8-27B beats Opus 4.6” can be true for a selected coding row and still be false as a general statement. In the Arena Code/WebDev snapshot dated 15 August 2026, Qwen3.8 Max and Grok 4.6 rank above Opus 4.6 high; Qwen3.8-27B is not listed in that snapshot. The Artificial Analysis/community view places Qwen3.8-27B around 52, which is remarkable for a 27B model, but it is not evidence that it wins every benchmark, every tool loop, or every customer workload.

Use the claim this way in a customer conversation: **“Qwen3.8-27B is a frontier-like local option for its size; validate it against your own holdout before replacing a hosted frontier model.”**

### Arena snapshots: preference, code, and agent behavior are different signals

Arena scores come from human or task-level comparisons and are best read as a snapshot of preference or practical performance. They are not interchangeable with Artificial Analysis’ composite index. The rows below are deliberately dated so the page does not imply that a leaderboard is static.

<figure><img src="https://487764224-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZ3ppnJjAH1qBNXEYnDPA%2Fuploads%2Fgit-blob-32b632d45f5853367abf0e545897b0ce55fa5b85%2Foh54-arena-code-snapshot.svg?alt=media" alt="Arena Code WebDev snapshot comparing six Office Hour 54 models with Opus 4.6"><figcaption><p>Arena Code/WebDev overall snapshot from 15 August 2026. Scores are preliminary and change as votes accumulate; <a href="https://arena.ai/leaderboard/code">open the live Arena leaderboard</a> before quoting a rank.</p></figcaption></figure>

Selected Arena Code/WebDev rows from that snapshot:

| Rank | Model                       | Arena score | Price shown by Arena                     |
| ---: | --------------------------- | ----------: | ---------------------------------------- |
|    3 | Qwen3.8 Max                 |   1667 ± 13 | $2 / $6 per 1M input/output tokens       |
|    5 | Grok 4.6 high               |   1631 ± 17 | $2 / $6 per 1M input/output tokens       |
|    8 | Gemini 3.7 Flash high       |   1587 ± 13 | $0.75 / $3.57 per 1M input/output tokens |
|   10 | DeepSeek V4 Pro high (0813) |   1584 ± 14 | $1.32 / $3.96 per 1M input/output tokens |
|   16 | Opus 4.6 high               |    1545 ± 6 | Provider price varies                    |
|   22 | Muse Spark 1.2 xHigh        |   1535 ± 14 | $1.25 / $4.25 per 1M input/output tokens |

The separate [Agent Arena leaderboard](https://arena.ai/leaderboard/agent) is a better lens for tool loops and steerability, but it does not contain every model in this release. In the 13 August 2026 snapshot, Qwen3.8 Max is #11 with 12.54% confirmed success and 0.11% tool hallucination, Gemini 3.7 Flash high is #21 with 9.85% confirmed success and 1.18% tool hallucination, and DeepSeek V4 Pro is #28 with 2.16% confirmed success and 0.35% tool hallucination. The absence of Grok 4.6, Muse Spark 1.2, or Qwen3.8-27B from that table means “not covered in this snapshot,” not “failed.”

### Qwen3.8-27B: a local frontier experiment on a practical budget

The most actionable customer story in this release is not a leaderboard screenshot. It is the ability to run a strong open model close to the data:

1. **Start with a quantized build.** Unsloth’s current guidance puts a 4-bit Qwen3.8-27B build in the roughly 17–19GB range. A 24GB GPU or a modern Apple-silicon system with sufficient unified memory gives safer headroom than a machine that only meets the weight-file minimum.
2. **Budget for the whole loop.** KV cache, context length, tool schemas, retrieval results, the runtime, and the operating system all compete for memory. “Fits in 17GB” does not mean “runs a 1M-token agent at full speed.”
3. **Use a bounded first demo.** Point the model at a private repository or document set, let it propose a patch or answer with citations, and require human approval before any external write. This demonstrates privacy and cost without pretending that local inference is automatically autonomous.
4. **Measure the customer’s outcome.** Record task success, review edits, latency, tokens, memory pressure, and failure recovery. Compare those numbers with the hosted model that the customer would otherwise buy.

<figure><img src="https://487764224-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZ3ppnJjAH1qBNXEYnDPA%2Fuploads%2Fgit-blob-d732257efdc4cba4dd8aa5c202a336e22ff97683%2Foh54-qwen-local-stack.svg?alt=media" alt="Diagram showing a private data workflow using Qwen3.8-27B locally with a quantized runtime and approval gate"><figcaption><p>Reference architecture for a local Qwen3.8-27B pilot. The memory figure is a quantized-weight estimate; measure context, throughput, and tool reliability on the target machine.</p></figcaption></figure>

An illustrative $2,000 workstation can therefore be a compelling pilot budget, but it is not a promise of a particular build, GPU, throughput, or total cost of ownership. Hardware prices, quantization formats, and context requirements change quickly; quote a current configuration only after measuring the customer’s workload.

### DeepSeek Harness: the new layer underneath the model

[DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) is a developer-preview harness, not a seventh model in the catalog. Its design treats the model, tools, context, sessions, and sandbox as replaceable plugins. The official [DeepSeek Harness site](https://deepseek.com/harness/) describes modes for standard work, coding, minimal interaction, and creator-style flows; the repository documents an append-only session log that supports replay, resume, fork, and search.

That separation is valuable for customers because it makes the experiment reproducible:

* **Model test:** keep the prompt, tools, and task fixed while comparing Gemini, Grok, Qwen, Muse, and DeepSeek.
* **Harness test:** keep the model fixed while comparing session memory, tool permissions, retries, sandboxing, and approval gates.
* **Operations test:** keep both fixed while measuring cost per successful task, time to recovery, trace completeness, and data residency.

DeepSeek Harness is still a preview and may change incompatibly. It should be treated as an interoperability and reproducibility idea to learn from, not as a guarantee that every AgenticFlow workflow or provider route will behave the same way.

### What people are trying in the open model ecosystem

Social posts are useful for finding experiments, but they are anecdotes rather than controlled benchmarks. These searches are intentionally linked so readers can inspect the original thread and date:

* **Local coding:** community posts around [Qwen3.8-27B](https://x.com/search?q=%22Qwen3.8-27B%22\&src=typed_query) focus on 4-bit quantization, 24GB-class machines, and private repository work. Treat “runs on my machine” as a starting point, then reproduce it with the customer’s context and tools.
* **Long-horizon coding:** [Muse Spark 1.2 discussions](https://x.com/search?q=%22Muse%20Spark%201.2%22\&src=typed_query) highlight persistent subagents, visual-to-code work, and long tool sequences. The interesting idea is the harness plus the model, not just the model name.
* **Voice and parallel agents:** [Grok 4.6 agent threads](https://x.com/search?q=%22Grok%204.6%22%20agent\&src=typed_query) show people trying voice prompts, parallel research, and coding loops. Verify what was actually executed before presenting a demo as an automation claim.
* **Reproducible sessions:** [DeepSeek Harness threads](https://x.com/search?q=%22DeepSeek%20Harness%22\&src=typed_query) are exploring replayable logs, plugins, and local sandboxes. This is a useful design pattern for auditability even when the underlying model changes.

### Customer value: turn a big model release into a smaller buying decision

For a customer, the release is valuable when it reduces risk or cycle time—not when it adds six names to a selector. A practical pilot can be designed in one afternoon:

| Step                     | Customer question                                                                 | Evidence to keep                                                  |
| ------------------------ | --------------------------------------------------------------------------------- | ----------------------------------------------------------------- |
| 1. Choose the lane       | Is the priority speed, frontier quality, local privacy, or long-horizon tool use? | One sentence naming the workload and data boundary                |
| 2. Build a holdout       | What five to ten real tasks represent the work?                                   | Inputs, expected outputs, and a review rubric                     |
| 3. Run two routes        | What changes between hosted and local execution?                                  | Quality, latency, token use, memory, and cost per successful task |
| 4. Add the safety gate   | Which actions require confirmation or a draft-before-send step?                   | Tool permissions, approval screenshots, and failure traces        |
| 5. Ship the smallest win | Can one workflow save time without claiming an external action occurred?          | Before/after time, acceptance rate, rollback path, and owner      |

The release’s safer delete confirmations, clearer MCP status, and better form labels make this evaluation easier to operate. The model expansion then gives the customer a controlled choice: fast multimodal work with Gemini, frontier hosted work with Grok or Qwen Max, harness-led long-horizon work with Muse, open-weight text/code work with DeepSeek, or a private Qwen3.8-27B pilot close to the data.

#### 2. Safer Delete Actions Across the Workspace

Deleting important workspace items now includes a clear confirmation step in more places.

Before removal, AgenticFlow asks you to confirm when deleting:

* A workflow
* A file or folder in AI Drive
* A connected MCP server
* A Workforce
* A saved connection

The confirmation identifies what will be removed, helping prevent accidental clicks while keeping the final decision with the user.

<figure><img src="https://487764224-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZ3ppnJjAH1qBNXEYnDPA%2Fuploads%2Fgit-blob-475807f7456b8d6a25dd60f381e9c9db494d9b34%2Foh54-mcp-delete-menu.png?alt=media" alt="MCP connection list showing the Delete action in an item menu"><figcaption><p>A connection item exposes Delete as an explicit action before the confirmation step.</p></figcaption></figure>

Dataset actions have also been improved so choosing **Delete** opens the expected confirmation instead of taking you to the dataset details page.

#### 3. Clearer MCP and Connection Guidance

MCP pages now give a more accurate picture of whether a server is ready to use.

If a connection needs attention or its tools are not currently available, AgenticFlow now shows clearer guidance about the next step and tool availability.

<figure><img src="https://487764224-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZ3ppnJjAH1qBNXEYnDPA%2Fuploads%2Fgit-blob-6a8dfbce6384ed20196d3bb343d7aa53aac4100e%2Foh54-mcp-connections.png?alt=media" alt="AgenticFlow MCP client connection list"><figcaption><p>The MCP client connection page groups saved servers in one place and keeps the available actions visible.</p></figcaption></figure>

Additional connection improvements include:

* Clearer connection-method and sign-in choices when adding a custom MCP server
* Better identification of connected servers and their available tools
* A clearer distinction between connection status and available actions
* More focused next-step actions on the OpenAI-compatible provider setup page

These changes make it easier to understand whether an integration is ready before attaching it to an agent or workflow.

#### 4. More Predictable Creation and Navigation

Several journeys now take users directly to the intended destination:

* **Library:** Selecting **Create resource** now opens the resource editor correctly, including when the create page is opened directly.
* **Marketplace:** Selecting **View all** for Multi-Agent templates keeps the Multi-Agent category active.
* **Welcome:** The signed-in **Deploy your agent** button now starts the Telegram setup when needed and continues into onboarding after setup.
* **Datasets:** Delete actions stay within the expected confirmation journey.
* **OpenAI-compatible providers:** The setup screen now presents only useful next actions.

These updates reduce dead ends and make common setup paths feel more consistent.

#### 5. Easier Agent and Workflow Building

Forms and controls across the builders now explain their purpose more clearly, including when used with a keyboard or screen reader.

Improvements cover:

* Workflow run inputs and required fields
* Workflow and agent schedule settings
* Workflow and agent variables
* Agent webhook settings
* Image generation options such as model, prompt, quality, aspect ratio, steps, and connection
* API key creation
* Workspace member invitations and project member roles
* AI Drive folder creation
* Bulk-run table creation
* Telegram bot setup

Visible labels now stay connected to the field they describe, and selection controls provide clearer context than the selected value alone. This makes complex forms easier to scan and reduces uncertainty when moving between fields.

<figure><img src="https://487764224-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2FZ3ppnJjAH1qBNXEYnDPA%2Fuploads%2Fgit-blob-c479c1c77a38fc53e2b4bbc8d30fa0f346e4fd6c%2Foh54-agent-editor-controls.png?alt=media" alt="AgenticFlow agent editor showing named capability and builder controls"><figcaption><p>The agent editor presents named capability, knowledge, memory, and chat-experience controls as scannable rows.</p></figcaption></figure>

#### 6. Clearer Sharing, Publishing, and Chat Controls

AgenticFlow now gives similar actions distinct, descriptive names so users can tell them apart more easily.

Updates include:

* Workflow sharing options that clearly identify what each switch controls
* Agent settings switches with clearer names and selected states
* Agent publish visibility that is easier to identify
* Platform-specific names for each **Configure** action on the publish screen
* Template **Duplicate** actions that include the template name
* Agent version and chat-history actions that are easier to recognize
* Agent sharing actions that work with both keyboard and mouse input
* Clearly identified message and send controls in agent chat
* More usable Workforce publishing and template-cloning choices

The create-app flow also now communicates which app type is selected, making it easier to review a choice before continuing.

#### 7. Better Keyboard and Assistive-Technology Support

This release includes a broad usability pass across common AgenticFlow screens. Buttons that previously relied only on an icon now have meaningful names, choice cards can be reached and selected from a keyboard, and switches communicate what they control.

The improvements are available across:

* Agents and agent chat
* Workflows and schedulers
* Workforce publishing and cloning
* Marketplace templates
* MCP and connection setup
* Projects and member invitations
* AI Drive and Datasets
* Image generation
* Pricing information
* Sharing and publishing

These changes do not alter the underlying workflow. They make the same features easier to understand and operate for more users.

***

### Model Updates Included

* Added Gemini 3.7 Flash
* Added Grok 4.6
* Added Muse Spark 1.2
* Added Qwen 3.8 Max
* Added DeepSeek V4 Pro 0813
* Added Qwen3.8-27B

### Improvements Included

* Added confirmation before deleting workflows, AI Drive items, MCP servers, Workforces, and connections
* Improved Library resource creation and direct access to the create page
* Improved MCP connection-status and tool-availability guidance
* Corrected Multi-Agent Marketplace navigation
* Improved the signed-in welcome and deploy journey
* Kept dataset deletion separate from detail-page navigation
* Added clearer labels across agent, workflow, workforce, scheduler, project, connection, and image-generation forms
* Improved keyboard access for sharing, publishing, category selection, and template cloning
* Made repeated actions easier to distinguish by including the relevant platform, template, or item name
* Improved chat message, send, history, version, and close controls
* Clarified selected states for app types, switches, and choice controls

***

### 🚀 6 Workflow Blueprint Prompts

Copy any prompt below into **Claude Code**, **Codex**, **Ishi**, or another AI coding agent with the **AgenticFlow CLI** installed to deploy a working workflow in minutes. These prompts were selected from six different AgenticFlow use cases and are ready to run as one-shot blueprint demos.

#### 1. Translation Quality Check Pack (`translation-quality-check-pack`)

```
Using the AgenticFlow CLI, deploy the translation-quality-check-pack blueprint. Use this Japanese source text:

"本契約は、顧客データの処理、保存、削除に関する責任分担を定めるものです。サービス提供者は、監査ログを90日間保持し、顧客の要求に応じてエクスポート可能にします。"

English translation: "This agreement defines each party's responsibilities for processing, storing, and deleting customer data. The service provider keeps audit logs for 90 days and can export them when the customer requests it."

Deploy, run once, show me the overall quality, meaning drift findings, terminology issues, suggested revision, and reviewer notes.

Leave the workflow deployed. Print the Web UI link.
One-line note: why this rung of the composition ladder?
```

***

#### 2. Add Sales CSV Invoice Pack workflow blueprint (`sales-csv-invoice-pack`)

```
Using the AgenticFlow CLI, deploy the sales-csv-invoice-pack blueprint. Use this CSV: customer,product,quantity,unit_price / Acme Co,Workflow Setup,1,1200 / Acme Co,Support Hours,8,150 / Brightline Inc,Training Seat,5,200. Invoice policy: "Calculate draft invoice totals by customer. Do not claim invoices were emailed, stored, paid, or approved." Deploy, run once, show me draft invoice summaries, validation issues, and next steps. Leave the workflow deployed. Print the Web UI link. One-line note: why this rung of the composition ladder?
```

***

#### 3. Convert website knowledge base workflow blueprint (`website-knowledge-base-pack`)

```
Using the AgenticFlow CLI, deploy the website-knowledge-base-pack blueprint. Use https://pixelml.com/ as the source URL.

Knowledge base policy: Create a RAG-ready knowledge base outline with source summary, sections, candidate chunks, metadata suggestions, unanswered questions, and QA checklist. Do not claim crawling beyond the supplied URL, vector database write, Supabase update, or external action occurred.

Deploy, run once, and show me the source summary, section outline, candidate chunks with metadata suggestions, unanswered questions, and QA checklist.

Leave the workflow deployed. Print the Web UI link.
One-line note: why this rung of the composition ladder?
```

***

#### 4. Company Docs QA Packet Workflow Blueprint (`company-docs-qa-pack`)

```
Using the AgenticFlow CLI, deploy the company-docs-qa-pack blueprint. Use these document excerpts:

Section: Expense Policy
Employees may expense economy airfare, hotel rooms up to USD 250 per night, and client meals up to USD 75 per person. Manager approval is required for exceptions.
---
Section: Security Review
New vendors must complete security review before production access. Required evidence includes SOC 2, DPA, subprocessor list, and SSO support.
---
Section: Onboarding
New hires receive laptop provisioning, HR orientation, and finance setup during week one.

Question: What evidence do we need before giving a new vendor production access, and who approves expense exceptions?

Deploy, run once, show me the direct answer, citation table, confidence, missing information, and suggested follow-up question.

Leave the workflow deployed. Print the Web UI link.
One-line note: why this rung of the composition ladder?
```

***

#### 5. Attendance Analytics Alert (`attendance-analytics-alert`)

```
Using the AgenticFlow CLI, deploy the attendance-analytics-alert blueprint. Use these attendance records:

"2026-06-04 | Alex Chen | Engineering | Present | Check-in 09:04 | Check-out 17:45
2026-06-04 | Mia Torres | Design | Late | Check-in 10:22 | Check-out 18:10
2026-06-04 | Dan Patel | Sales | Absent | Reason: no notice
2026-06-04 | Sarah Kim | Support | Present | Check-in 08:55 | Check-out 17:05"

Attendance policy: "Late after 09:30. Absence with no notice is critical and should alert the manager. More than one late/absent record in a department should be flagged for HR review."

Deploy, run once, show me the manager alert message, HR summary, exception table, and follow-up checklist.

Leave the workflow deployed. Print the Web UI link.
One-line note: why this rung of the composition ladder?
```

***

#### 6. Convert phase blog production workflow blueprint (`phase-blog-production-pack`)

```
Using the AgenticFlow CLI, deploy the phase-blog-production-pack blueprint. Use this blog brief:

"Topic: How AI workflow automation reduces manual operations work | Audience: operations leaders at B2B SaaS companies | Goal: generate qualified demo interest | Primary keyword: AI workflow automation | Tone: practical and executive-friendly | Target length: 1,200 words"

Production rules: Create a phase-based production packet with strategy, outline, draft plan, review checklist, SEO notes, and publishing handoff. Do not claim CMS publishing, Google Docs updates, agent delegation, or external action occurred.

Deploy, run once, and show me the strategy summary, phase table, outline, QA checklist, SEO notes, and publishing handoff.

Leave the workflow deployed. Print the Web UI link.
One-line note: why this rung of the composition ladder?
```

### Get Started

* **Try AgenticFlow free:** <https://agenticflow.ai/>
* **Latest changelog:** <https://docs.agenticflow.ai/changelog>
* **Join our Discord:** <https://qra.ai/discord>
* **Need help?** Email <support@agenticflow.ai> with the page you were using and what you expected to happen.
