Ask a model whether your long is a good idea and it will tell you it is. Open a fresh chat, describe the same market with the opposite position and it will tell you that this, too, is a good idea. This is because the model is engineered to agree with you. When LLMs are being developed, humans score their answers and intuitively pick the ones where the model is more agreeable, turning them into a "yes-man".
This article consists of three parts: what the machine is reliably good and bad at, the baseline of a prompt that produces work, and then the patterns that exist purely to fight the failures. This info is useful regardless of whether you trade full-systematic, discretionary, or with a mix of both.
Disclosure: Purple Technology, which publishes Hello Purple, builds products in agentic finance, so read this with that interest in mind.
What it is good at, and what it is bad at
Broad, heavily documented territory is where it shines. Ask what basis risk is, why a long position in a futures curve in contango bleeds on every roll, how volatility drag eats the compounded return of a leveraged position, what a triple witching event is and why it matters. There are thousands of explanations across the internet, so answers generally hold up and need the least scepticism. The same is true for reorganising material you give it: pull every strike and expiry out of this transcript into a list, turn this broker statement into rows, find every mention of inventory in this filing and quote the sentence.
On the other end of the spectrum, specifics are its biggest weakness. Exact figures, dates, tick values, contract multipliers, margin requirements, settlement conventions, all data that changes all the time, and which the model can respond to with a plausible looking, yet fake number. Or, it could give you data from the end of its training cut-off, thinking that that's today's data. These seem like small things, but they can screw over your actual risk with multiples the size of what you intended.
Calculation is another one of these weak points, and feels counterintuitive. We'd think a computer would be amazing at math, especially seeing how well the models can explain mathematical concepts. Yet, we cannot trust them with any complicated calculations, as it's predicting tokens rather than computing. The way to solve this is to either do the calculations yourself, or code it.
Finally, there's one more key issue, which lies in recognising the wrong answers. As humans, we have intuitively learned that confident answers mean that there's a high likelihood that the response is right. But a machine doesn't know confidence, so the numbers it just made up will be described with the same confidence as the ones it just got from a highly respected source.
Three biases to watch out for
It agrees with what you already think. Modern chat models are tuned on responses that human raters preferred, and a 2023 study by Sharma and colleagues, published at ICLR 2024, found that a response matching the user's views is more likely to be preferred, and that both humans and preference models pick convincingly written sycophantic answers over correct ones a non-negligible fraction of the time. A Stanford evaluation found sycophantic behaviour (a model's inclination toward agreeing with the user's stated view over the correct answer, produced by preference tuning) in 58.19% of cases across three frontier models, measured on mathematics and medical-advice datasets rather than markets, and a 2026 benchmark on agentic financial tasks found that most models fail once user preference information contradicts the correct answer. For a trader this is close to the worst available failure mode, because you almost never arrive neutral. You arrive with a position on, or one you want to open, and what you want is permission. A model that mirrors your framing is a confirmation-bias engine that feels like outside validation precisely because it came from outside your head.
It flatters the work. Ask it to critique an analysis it produced earlier in the same conversation and it goes easy on itself. Everything already in the thread is context it treats as established, so the critique anchors to the thing it is meant to attack, and you get a few softened quibbles that leave the conclusion standing. Evaluation research documents a related bias: in tests on text summaries, models scored their own outputs higher than those of other models or humans, while human annotators rated them as equal in quality.
It will overfit anything you point it at. Ask an agent to improve a strategy's backtest and it will do exactly that: go into the historical results, find every cluster of losing trades, and propose filters that remove them. Skip these hours, avoid Mondays, no entries when this indicator reads above that threshold. Every suggestion improves the backtest, because every suggestion was reverse-engineered from the backtest. The agent is simply doing what you told it to. You asked for better historical performance, and mining the past for exceptions is the most direct route to one. The safe default is to treat any AI-suggested improvement to a backtest as overfitting until proven otherwise.
How to prompt the proper way
Instead of going constantly back and forth over a 2-hour timespan, take more time to craft your prompt. Most disappointing answers that are not explained by the weaknesses above trace back to the prompt. Use the four slots below as a checklist before you send it.
Context is everything it cannot know: no live prices, no view of your platform, no idea what today is unless told. "Is selling this straddle a good idea" gets you a textbook answer. The same question carrying the 30-day implied vol, the 20-day realised, days to earnings, the implied move and the average absolute earnings move over the last 12 quarters gets you arithmetic on real inputs. The biggest single upgrade available to most people is pasting more of what they are already looking at.
Task is a verb with a checkable result. Vague verbs invite essays: analyse, discuss, review, consider. Specific verbs invite work: list, rank, calculate, compare, identify. When you catch yourself typing "thoughts on this?", stop and ask what you would do with a good answer, then request that.
Constraints say what to skip and what to assume. One is worth more than it looks: if an input you need is missing, say so instead of estimating it. That makes the honest response an allowed completion rather than an apparent failure, and it measurably changes behaviour. Anthropic's guidance on reducing hallucinations lists explicit permission to say "I don't know" as a technique that can drastically reduce false information.
Format is the shape of the answer, and it does more than tidy the output. Ask for one row per scenario with columns for trigger, main risk and invalidation level, and every cell has to be filled. Prose can sound balanced while committing to nothing.
Three habits from Anthropic's own guidance are added on top of that. Say why: giving the reason behind an instruction works better than the bare instruction, because the model generalises from the explanation, and the trading version is direct, since "these figures feed a position size, so if a number is not in what I gave you, do not supply one" beats "do not make up numbers". Show rather than describe: their docs put examples among the most reliable ways to steer format and structure. When your prompt includes a lot of text, such as several documents, reports or transcripts, you should paste that material first and put your actual question and instructions at the end, which they report can improve response quality by up to 30% on complex multi-document inputs.
Prompting against the weaknesses
Blind the analyst. Never reveal which side you are on before asking for analysis. Instead of "I'm long crude from 71.50, should I hold through the inventory report", write "evaluate the case for being long crude here and the case for being flat, given tomorrow's inventory report", then paste the data. There is now no side to flatter, and you apply the result to your own position yourself.
Attribute it to a stranger. When the direction is impossible to hide, hand it to someone else. "A trader I follow proposes shorting this into earnings on the following reasoning. Assess the argument." Models critique a third party's reasoning far more freely than they critique yours, because there is no user to please on the other side of it. A 2026 study of seven models found that between 27.7% and 46.3% of the critical judgments they gave about an anonymous person's actions disappeared when the same actions were presented as the user's own, and a 2025 study found that a third-person framing cut sycophancy by up to 63.8% in debate scenarios. Presented as your plan, the same reasoning gets handled with kid gloves. The asymmetry is absurd and completely exploitable.
Demand equal effort on both sides. Ask for the strongest bull case and the strongest bear case at roughly equal length, then which one the data you supplied supports better. Without the equal-effort clause you get four paragraphs for the side your phrasing hinted at and one dutiful paragraph for the other.
Ask to be talked out of it, and set the stakes. Once you hold a view, the useful question is not whether you are right but what the best argument against you is. Argue against this trade as if you were paid to talk me out of it, do not soften it, I want the strongest opposing case rather than a balanced summary. This works because it redefines pleasing you as producing a hard counterargument, so the eagerness finally runs in your favour.
Establish the base rate before the setup. Ask about the general category first: how often breakouts from multi-month ranges follow through versus fail, what typically happens to implied vol in the week after earnings. Then ask how your case differs from the average. Present your specific trade first and everything after it is coloured by it; establish the anchor first and your trade has to argue against it. Treat any precise-sounding figure it gives you as unverified, because this is the territory where it fabricates.
Make it label its own confidence. For every factual claim, high, medium or low, where low means it could not verify the claim from what you gave it. The labels are not calibrated probabilities and you should not read a "high" as 90%. What they do is separate claims it is pattern-matching confidently from claims it is generating to fill space, and you will find the low labels cluster on the exact figures, dates and named facts where fabrication lives. That tells you where to spend your checking.
Split the writer from the critic. Generate in one conversation. Open a fresh one, paste the output with no history attached, and frame it as another analyst's work: find the weakest points, the unstated assumptions, the missing scenarios, and any place the reasoning breaks if one stated input is slightly wrong. The fresh context has no investment in the conclusion and nothing earlier to stay consistent with, and the difference in critique quality is not subtle.
Ask for falsification, never optimisation. This is the counter-move to the overfitting bias, and it is one word's difference in the request. Not "improve this strategy's performance" but "find reasons this strategy should not work", "identify which of these rules lack a structural justification", "check this code for look-ahead bias", "list the ways this rule could look profitable in a backtest while being untradeable live". Agents struggle strongly with improving backtests without overfitting, even when being prompted explicitly to do so, as they'll simply conjure up a plausible explanation to any pattern in the data they see. In a 2026 test, an evaluation that corrected for the size of the search and for look-ahead bias rejected every strategy that two frontier models had discovered, with search budgets of up to 100 candidates. Reversing their task however, stress-testing and pulling a strategy apart, is one of their strong suits, and they bring the objectivity of someone with nothing invested in the work.
Make claims checkable rather than plausible. Anything factual gets a verbatim quote from the source you supplied, with its location, and an explicit "no direct support in the document" where there is none. A fabricated paraphrase costs you a reading session to catch. A fabricated quote costs you one search. If a quote you spot-check fails, the whole summary is suspect rather than that one line, because the model was in generating mode for part of the task and you cannot know which part.
Test the prompt, not the conversation. When an answer is bad, the instinct is to reply "no, not like that" and steer until the output is acceptable. That fix evaporates when the chat ends. Work out which of the four slots failed, rewrite the prompt, and rerun it clean, so the improvement is the thing you use next time. Then test the rewritten version the way you test anything else: on a case where you already know the answer. Before trusting a journal-review prompt, run it on a month you have already gone through by hand and see whether it finds what you know is in there. Change one thing at a time.
The thing to try this week is the cheapest one on the list. Take a position you currently hold, write the naive prompt that reveals your side, write the blinded version with both cases at equal length and stated confidence per claim, and run them in separate chats.
This article is information and commentary. It is not investment advice, not a personal recommendation and not investment research, it takes no account of your circumstances, and it must not be relied on as a reason to trade. Any figure shown is historical or simulated and is not an indication of future results.
Get the next piece the day it is published.
One email when a new report lands, nothing else.
Exact figures, sources, steps and terms behind this piece, written for the coding agent you point at it.
Quote freely with attribution: name the author, Hello Purple and the canonical URL, with the publication date, all of which are in the Markdown twin’s front matter. The twin is the reference text of the piece.
Read the appendix
Facts and figures
- SycEval (arXiv 2502.08177, Stanford) observed sycophantic behaviour in 58.19% of cases across ChatGPT-4o, Claude Sonnet and Gemini 1.5 Pro, measured on the AMPS mathematics and MedQuad medical-advice datasets. Highest rate Gemini at 62.47%, lowest ChatGPT at 56.71%.
- "The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications" (arXiv 2604.24668, 2026) reports that most tested models fail when given user preference information contradicting the reference answer.
- "Towards Understanding Sycophancy in Language Models" (Sharma et al., arXiv 2310.13548, ICLR 2024) found in human preference data that a response matching a user's views is more likely to be preferred, and that both humans and preference models prefer convincingly written sycophantic responses over correct ones a non-negligible fraction of the time.
- Anthropic's prompting guidance states that placing queries at the end of a prompt, below long documents, can improve response quality by up to 30% on complex multi-document inputs, and recommends 3 to 5 few-shot examples.
- "Affective Context Amplifies Sycophancy in LLM Responses" (arXiv 2608.21242, August 2026) tested seven LLMs and found that 27.7% (Claude) to 46.3% (Llama) of critical judgments made about an anonymous person's actions were withheld when the same content was presented as the user's own.
- "Measuring Sycophancy of Language Models in Multi-turn Dialogues" (Hong et al., arXiv 2505.23840, Findings of EMNLP 2025) reports that a third-person perspective reduced sycophancy by up to 63.8% in the debate scenario.
- "What survives honest evaluation? Leakage-safe, search-aware assessment of LLM-driven trading strategy discovery" (arXiv 2608.27734) rejected every LLM-discovered strategy across two frontier models, search budgets up to 100 candidates and five repeated runs.
- "LLM Evaluators Recognize and Favor Their Own Generations" (Panickssery et al., NeurIPS 2024, arXiv 2404.13076) found that LLM evaluators score their own outputs higher than those of other models or humans, while human annotators consider them of equal quality.
- Anthropic's guidance on reducing hallucinations lists explicit permission to say "I don't know" as a technique that can drastically reduce false information.
Steps to reproduce
Reproducing the agreement effect described at the top of this piece:
- Choose a position currently held, or a market where a view is already formed.
- In a fresh conversation, state the position held and ask whether it is a good idea. Record the conclusion.
- In a second fresh conversation with no history, state the opposite position on the same market, with the same supporting data, and ask the same question. Record the conclusion.
- In a third fresh conversation, supply the data only, disclose no position, and request the case for each side at roughly equal length, a stated conclusion and a confidence label per factual claim.
- Compare the three outputs. Steps 2 and 3 will typically each endorse the position stated. Step 4 is the one to act on.
- Repeat on a closed trade whose outcome is already known, as a check on the blinded version against a settled case.
Limits
A chat assistant without tools has no live market data and no knowledge past its training cutoff. Accuracy is lowest on specific figures, dates, contract specifications and venue fine print, and highest on broad well-documented concepts and on restructuring supplied text. Multi-step arithmetic is unreliable without code execution. Response tone does not vary with correctness. A request to improve a backtest returns filters derived from that backtest.
Sources and links
- https://arxiv.org/abs/2502.08177
- https://arxiv.org/abs/2604.24668
- https://arxiv.org/abs/2310.13548
- https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices
- https://arxiv.org/abs/2608.21242
- https://arxiv.org/abs/2505.23840
- https://arxiv.org/abs/2608.27734
- https://arxiv.org/abs/2404.13076
- https://docs.anthropic.com/en/docs/test-and-evaluate/strengthen-guardrails/reduce-hallucinations
Terms
- Sycophancy: a model's tilt toward agreeing with the user's stated view over the correct answer, produced by preference tuning.
- Overfitting: tuning rules so closely to historical data that they describe past noise rather than market behaviour.
- Look-ahead bias: a backtest using information that was not available at the moment the decision was made.
- Training cutoff: the date after which a model holds no knowledge from training.
- Few-shot prompting: steering output by supplying worked examples of the input and output wanted.
About this piece
Hello Purple, the-map track, 2026. Purple builds products in agentic finance.
Financial Market Analyst at Purple Technology. Holds a bachelor's degree in Finance and Insurance from HOGent, graduated with honours, and previously worked as an investment analyst at a family office. Spent two years in an investment club, analysing stocks and preparing pitches for the other members. Has traded his own account for five years across risk-premia harvesting, order flow trading and discretionary momentum swing setups, and built his own strategies on well-known concepts like open drives, inside days and open gaps. Writes here about AI in finance in practice, with a focus on trading and investing.