A successful tool call leaves an unanswered question
Connecting an AI assistant to a useful service changes what it can know. It also creates another place where a claim can go wrong. The service returns information and an interpretation. The assistant decides what to carry into its answer. What survives that handoff?
We studied this using Draconic, the market-information service we build. The connected assistant received intraday price sequences that ordinary web research often did not retrieve. Those sequences let it answer questions that would otherwise remain unresolved.
But the final answer was not a simple copy of the service response. The assistant omitted several unsupported interpretations from the service. Other mistakes reached the reader. Useful evidence and faulty claims could travel through the same connection.
That distinction matters beyond trading. Any assistant that explains information from another system makes choices about what to preserve, qualify or leave out. We wanted to inspect those choices at the point where the service's output becomes the user's answer.
One assistant, two ways of getting information
We collected sixteen answers from Astra 6 across four instruments: spot gold, the NIFTY 50 index, ICICI Bank shares and Reliance Industries shares. Each instrument had an initial question and a follow-up. Both questions were answered with ordinary web research and with web research plus access to Draconic.
The assistant and its general instructions stayed the same. Its available information changed. Draconic supplied additional data, calculated measures and its own generated analysis. Each condition retained its own conversation history.
The questions asked for explanations across different time intervals, evidence of broad participation, or confirmation of a move through a price level. The follow-ups asked what had changed. Gold observations came from September 8, 2026; the Indian instruments came from September 9.
This was a comparison of access to the whole service. We did not give both conditions identical raw data, so we cannot attribute a useful answer to a particular Draconic calculation. We also did not place trades or measure returns. The full eight-question inventory, covering both conditions, appears in the methods.
The service and the assistant deserve separate inspection
The clearest benefit was access to the missing sequence of events. In the initial ICICI Bank question, ordinary research found a recent quote but could not establish the requested five-minute price sequence or hourly structure. The answer acknowledged insufficient evidence.
The connected assistant could inspect that sequence. A small green candle meant the interval closed above its opening price. Yet its high, low and close were all below the preceding candle's values. The answer correctly explained why one green candle did not establish a reversal. It also disclosed that the hourly background came from the previous day.
The extra information therefore answered a real part of the question. That does not make every interpretation supplied with it trustworthy. Some direct Draconic responses attributed price action to institutional intent, absorption or the causes of buying and selling without sufficient support in the captured evidence. Astra omitted several of those claims from its final answers.
Reliance showed the other outcome. The connected answer called ₹1,280 the morning low. The retained opening candle had reached ₹1,277.20. The ₹1,280 observation was a later pullback low. The incorrect label appeared again in the follow-up, even as the assistant usefully updated its interpretation of newer prices.
These observations do not identify why the error persisted. Conversation history and later service output can both contain earlier claims. What we can establish is that the final answer retained the wrong label. Reviewing only the service would miss the assistant's filtering; reviewing only the final answer would hide what happened between the two.
Reliance · 9 September 2026 · A recorded price became the wrong kind of reference.
- 09:15 low
- ₹1,277.20
- 09:25 pullback low
- ₹1,280.00
The later low was labelled the morning low.
The same morning-low label appeared in the answer.
Paraphrased from retained records. The numeric level was real; its label was incorrect. This traces an observed handoff, not the model’s hidden reasoning or the sole cause of the error.
Inspect the evidence
Both source captures contain the opening-candle low of ₹1,277.20. ₹1,280 is a later pullback low. The initial and follow-up assistant answers both describe ₹1,280 as the morning low.
The distinction changes what a later break of that level would establish. The study did not measure an order or trading loss caused by this error.
A condition can be partly observed and still remain unmet
The ICICI follow-up showed why a precise condition matters. The earlier answer required a completed close below ₹1,386.40 with follow-through. Later prices briefly moved below that level, but none of the four reported five-minute closes was below it.
The assistant distinguished a temporary breach during an interval from a closing break. It described a stronger rebound without declaring that the hourly structure had reversed. Ordinary research still lacked the newer sequence needed to make that assessment.
Gold supplied a different test. The initial condition required a fifteen-minute close below $4,382.12, followed by an unsuccessful attempt to recover the $4,382–$4,384 area. Those are two separate requirements: a closing break, then a failed recovery.
The follow-up supplied four new fifteen-minute closes and an hourly close below the earlier support area. It also stated that a direct retest from below was not demonstrated. The closing requirement was supported. The full condition remained unestablished.
For an assistant builder, the useful check is straightforward: compare every requirement in the original statement with the new observations. A fluent verdict such as “confirmed” can conceal a missing step. These examples show what to inspect; they do not prove that a particular checking procedure prevents the error.
Gold · The earlier confirmation condition required both observations.
A close below the level
New candle closes were below the earlier support area.
A failed reclaim
A direct retest from below was not demonstrated.
The closing component was supported. The full condition was not established.
Inspect the actual condition
The initial answer required a fifteen-minute close below $4,382.12 followed by a failed reclaim of the $4,382–$4,384 area. The follow-up provided later closes below that area but explicitly said the direct retest was not demonstrated.
This is a historical source-based illustration of condition checking. It is not a trading recommendation or proof that a proposed review method improves model performance.
Omitted evidence needs a comparability check too
The NIFTY follow-up exposed a harder problem than a wrong price label. The connected answer described a renewed index recovery alongside worsening participation. The number of advancing instruments had fallen from 54 to 50.
Two other supplied counts increased. Instruments above their average traded price, weighted by volume, rose from 57 to 74. Instruments above a moving average that gives more weight to recent prices rose from 46 to 70. The follow-up did not discuss those increases.
That omission is observable, but the meaning of the omitted counts is uncertain. These measures ask different questions. We did not independently establish that they covered comparable populations or calculation bases. Some sector and aggregate counts also did not reconcile.
We therefore cannot use the increases to declare broad support for the recovery. A stronger answer would explain which measures changed and where comparison remained unresolved. Checking an assistant's use of evidence includes checking whether apparently conflicting numbers can fairly be compared in the first place.
Keep the evidence beside the claim
These cases suggest a small review worksheet for assistants that use external services. For each important claim, retain the supporting observation, its timestamp and the wording delivered to the user. Then inspect what the assistant preserved, corrected, qualified or omitted.
For a follow-up, include the inherited claims in that review. A response can incorporate new information and still repeat an incorrect factual label. For a condition, record its requirements separately and mark which are observed, contradicted or still unverified.
This can be inspected at the service interface without publishing proprietary formulas. It also keeps the review fair: an assistant that admits it lacks a price sequence has a different limitation from one that invents or mislabels the sequence. We propose the worksheet as an evaluation method. Its effectiveness as a production safeguard remains untested.
What the evidence establishes
This is a small, vendor-authored study with AI-assisted assessment. The four situations were selected, follow-ups depend on initial answers, and the Indian instruments share a session. The source captures were not independently verified against exchange records. Some hourly information was dated.
We include all sixteen primary answers in the outcome inventory. An older 42-answer experiment remains separate development evidence. Different questions, dates and configurations prevent treating the combined material as 58 independent observations.
In these cases, additional market evidence made more specific answers possible. It did not guarantee that labels, conditions or qualifications were preserved. A successful tool response and a trustworthy final answer are separate outcomes. The distinction becomes visible when the source observations, service output and final answer are examined together.
Methods & evidence
The exact questions, timing, review limits and complete primary comparison inventory.
Authorship and assessment
This exploratory study was conducted within the Draconic project. The author has a commercial interest in the service. AI assistants helped with implementation, source review and synthesis. There was no independent human grading or peer review. Design consultations with other AI models are not independent validation of the results.
Separate source captures were saved before answer assessment. Project AI reviewers examined those sources before reading the corresponding answers, and the primary assistant reviewed their findings. This checks consistency with retained source data. It is not independent exchange verification or a blinded human preference study. External web claims were not exhaustively reverified.
Assessment criteria were developed and applied within the project. They were not registered before the experiment. The review asks what evidence each answer supplied, whether it used that evidence correctly, which misleading claims or omissions were found, and how it handled changed and unchanged observations. No comprehensive error rate, statistical significance claim or preference percentage is reported.
Configuration and observation times
Both conditions used gpt-6-astra at high reasoning effort through the OpenAI application programming interface. Both could research the web. One could additionally call Draconic. General instructions requested a concise read, evidence, what would change the assessment, and limitations with source times. They required separation of facts from interpretation and prohibited invented data or substitution of another instrument. The tool response was evidence, not instructions.
The evaluation software allowed at most three outer request rounds, an output limit of 8,192 tokens per round, and one Draconic analysis call per answer. Cases fixed the instrument, market and routing across time intervals. The study therefore does not evaluate autonomous instrument selection, installation, tool discovery or chart rendering. The ordinary-research condition did not receive the separate reference captures.
Draconic returned generated analysis with information about the available coverage and context. Its internal generation added computation beyond the final assistant's own request. The two conditions intentionally had unequal inputs. This does not isolate the contribution of proprietary measures, prompts or generated analysis.
All request times below are in Coordinated Universal Time, in 2026. A request time differs from a source observation time and from answer completion.
| Instrument | Initial request | Follow-up request |
|---|---|---|
| Spot gold | September 8, 18:46:03 | September 8, 19:50:36 |
| NIFTY 50 | September 9, 05:13:58 | September 9, 05:36:54 |
| ICICI Bank | September 9, 05:16:34 | September 9, 05:39:40 |
| Reliance Industries | September 9, 05:19:15 | September 9, 05:42:25 |
Each condition retained its own conversation history. Follow-ups were eligible at least twenty minutes after the initial pair completed. Gold's delay was longer. The last Reliance pair ran sequentially to preserve spending headroom; other pairs ran concurrently. Retrieval and completion times were not identical. Older hourly context remained available with its date, even when newer short-interval data arrived. The study does not claim perfectly synchronized market snapshots.
All sixteen primary answers
Each row represents two final answers to the same question, one per condition. These eight rows account for all sixteen primary answers. All completed. No complete primary pair was repeated to improve its outcome. Original answers remain unchanged; corrections are annotations. These summaries are qualitative observations, not grades or a win rate.
| Question | Ordinary web research | With Draconic | Qualification |
|---|---|---|---|
| Gold, initial | It found a broad hourly bearish reading but could not verify the shorter sequence. | It compared shorter candles with hourly structure and identified rejection near a low. | Source truth was not independently verified. |
| Gold, follow-up | It found newer bearish public evidence but still lacked the shorter sequence. | It identified later closes below the earlier support area. | The separate failed-reclaim requirement remained unverified. |
| NIFTY, initial | It described uneven sector weakness. | It supplied the intraday rebound and participation detail for its monitored universe. | Population definitions and some timestamp interpretations remained problematic. |
| NIFTY, follow-up | It distinguished newly retrieved old reports from new observations. | It identified a renewed recovery alongside weaker advance and decline counts. | It omitted improving technical counts, whose comparability remained uncertain. |
| ICICI, initial | It retrieved a recent quote but could not verify the required candle structure. | It distinguished a green candle from reversal of the declining sequence. | Hourly evidence came from the previous day. |
| ICICI, follow-up | It had no newer verified sequence. | It identified a brief undercut followed by recovery without the required closing break. | The hourly dataset remained unchanged. |
| Reliance, initial | It could not establish the required hourly level and candle test. | It explained a failed recovery through ₹1,288. | It incorrectly called ₹1,280 the morning low. |
| Reliance, follow-up | It found newer snapshots but could not establish the candle-level change. | It revised its interpretation after closes above ₹1,288. | The morning-low error persisted. |
Exact external question templates
These are the external evaluation questions supplied to the host assistant, not Draconic's production system prompt. Each placeholder was replaced with the actual instrument or request timestamp.
Gold: “Analyse XAUUSD spot gold in US dollars, as of {observation_time_utc}. Compare its one-hour structure with fifteen-minute momentum. Does the shorter-term move strengthen, weaken or leave the broader interpretation unchanged? Identify the most important supporting and contradictory observations, and what would make you reconsider. Do not assume the timeframes disagree.”
NIFTY: “Analyse the NIFTY 50 NSE spot index, as of {observation_time_utc}. Does its latest directional move have broad participation, or is the evidence concentrated in a few sectors or stocks? Explain what the index structure and available participation evidence support, what most challenges that interpretation, and what observation would change it. Distinguish your covered universe from the whole exchange. If using options evidence, identify the actual expiry and limit any conclusions to that contract.”
ICICI: “Analyse ICICI Bank's NSE cash shares, as of {observation_time_utc}. Compare the one-hour structure with the latest five-minute move. Is the shorter move evidence of a broader structural change, a counter-move within the existing structure, or is the evidence insufficient? Explain the decisive observations and the strongest contradiction to your assessment. Distinguish a completed candle from one still forming.”
Reliance: “Analyse Reliance Industries' NSE cash shares, as of {observation_time_utc}. Assess whether the latest five-minute price action around a relevant confirmed hourly structure level supports a breakout or continuation interpretation, or whether confirmation is still missing. Identify the level and evidence behind that judgement, including available momentum and volume. If there is no identifiable breakout attempt, say so rather than inventing one. What would invalidate your assessment?”
Follow-up: “Since your previous assessment of {instrument}, what has genuinely changed in the underlying market evidence, what remains intact, and does the new evidence change the interpretation? Distinguish a new observation from a newly mentioned old observation. If the data has not changed, say so. Observation request time: {observation_time_utc}.”
Source details behind the examples
Reliance's September 9 opening bar at 09:15 India Standard Time recorded a low of ₹1,277.20. The 09:25 bar recorded ₹1,280, a later pullback low. Initially, two five-minute candles traded above ₹1,288 and closed below it. Two new follow-up candles closed above ₹1,288. The assistant revised its interpretation to a tentative recovery without claiming a confirmed break above the earlier ₹1,289.70 high. The incorrect morning-low label remained.
ICICI's four follow-up intervals are shown below in India Standard Time on September 9. Their labels use the source's apparent bar-start convention. The capture did not include an explicit exchange-finality flag. None of the four closes was below ₹1,386.40.
| Five-minute interval | Low | Close |
|---|---|---|
| 10:45–10:50 | ₹1,387.50 | ₹1,387.50 |
| 10:50–10:55 | ₹1,385.70 | ₹1,386.40 |
| 10:55–11:00 | ₹1,385.50 | ₹1,387.10 |
| 11:00–11:05 | ₹1,387.10 | ₹1,389.90 |
NIFTY source captures were taken on September 9 at 10:43:58 and 11:06:54 India Standard Time. Advancing-instrument counts changed from 54 to 50. Counts above volume-weighted average price changed from 57 to 74. Counts above a 21-period exponential moving average changed from 46 to 70. These are distinct comparisons, with unresolved population and calculation comparability. Neither answer established a complete census of NIFTY constituents or the whole exchange.
Gold's initial condition combined a fifteen-minute close below $4,382.12 with a failed reclaim of $4,382–$4,384. The follow-up reported four new fifteen-minute closes and an hourly close below the earlier support area, while explicitly noting that a direct retest from below was not demonstrated.
The earlier development experiment is separate
An earlier experiment retained 42 selected answers across six task families and three host models: Astra 6, Fable 5.1 and Gemini 3.8 Flash. It covered NIFTY interpretation and follow-up, ICICI comparisons across time intervals, sector participation, unavailable options history, contract clarification and risk arithmetic. Indian-market observations were after close.
That work exposed errors in time-window labels, percentage baselines, options expiry and unsupported interpretations of order-book or dealer behaviour. Some cases found ordinary research adequate or no added production evidence from Draconic. Arithmetic and ambiguous-contract controls did not establish a production-data advantage.
The retained inventory contains 51 original and recovery task records. Eight original records stopped before provider execution because of concurrent budget reservations. There were 43 completed records and 42 final selected answers. Selection used the original unless a recovery existed for the same case, turn, model and condition. One completed original was superseded to retain a matched repeated follow-up; it remains preserved. Preflights are separate. Internal generation also included served-model fallbacks.
Different cases, timings and configurations prevent pooling the two studies into one score. They also prevent treating the primary results as a controlled demonstration that intervening fixes improved performance.
Recorded costs and their limits
Primary host-model token estimates totalled $10.32. A conservative spending tracker finished at $19.52: those estimates plus $8 reserved for internal generation and $1.20 reserved for search. The tracker is not an invoice or a measurement of actual internal generation cost. Visual work, search charges, cache adjustments, taxes and currency effects are not fully reconciled.
The earlier pilot, preflights and recovery recorded $24.02 in token estimates, including matched internal analysis usage and excluding other charges. Separate Fable design consultations recorded estimates of $0.50 and $0.37. A final-draft Fable consultation cost approximately $0.48 in token estimates. These amounts are outside the primary study's $19.52 tracker. Exact unrounded estimates remain in private receipts.
No further benchmark calls were made to prepare this publication draft. Drafting and review used existing project AI assistance, which is not represented as zero-cost computation.
Disclosure and limits of the comparison
This report includes external question templates, high-level configuration, assessment criteria, all primary outcome summaries and reviewed examples. Production system prompts, proprietary formulas and weights, complete internal feature definitions and payloads, customer records and bulk raw feeds are excluded.
Readers can assess the examples against the stated evidence. This does not reproduce the full proprietary engine or guarantee an identical historical response from changing data and model providers. An illustrative interface is not the production schema.
The connected condition added information access, calculated measures and generated analysis together. The comparison does not isolate the contribution of any one of these additions.