TECHUB NEWS Hong Kong Report | Reporter Alma Li | September 2026
Introduction:
As large models make strategy generation, information processing, and historical backtesting faster, how much of a seemingly beautiful result comes from truly independent, sufficient, and credible evidence? Bernard zkBernard, product manager at HighBlock Limited and independent developer, is developing the "Evidence Deflator" to address validation issues that are often overlooked in AI research investment. Before a strategy draws conclusions, he aims to audit model knowledge coverage, backtest periods, parameter searches, and statistical thresholds to judge whether this backtest truly has enough resolution.
In discussions of the AI research investment elite incubation program, Bernard zkBernard proposed an intuitive analogy: letting a large model validate a strategy using historical intervals already covered by its training data is akin to having a student take a test for which they already know the answers.
Core Viewpoint:
"Scores may be high, but what they test may not be ability; it could just be memory," he said.
The reason for choosing to approach AI research investment from the "validation layer": The market is not lacking in tools that use AI to generate opinions, filter assets, and build strategies; what is truly scarce is another ability: when a model gives a beautiful backtest, how does the researcher distinguish which elements are independent evidence and which are mere reutilization of known history by the model or researcher?
Currently, Bernard works at HighBlock Limited, focusing on product design and industry research for cryptocurrency institutional business within the Hong Kong compliance framework, covering wealth management, RWA, market-making, custody, OTC, AI, and other areas. Outside of work, he also engages in independent development and open-source research, focusing on the independence of AI judgments and the protocol interoperability between agent tools. Both paths ultimately converge on the same question: AI can generate conclusions faster, but whether those conclusions can withstand validation cannot be taken for granted.
A backtest can be an "open-book exam"
TECHUB NEWS: Many people consider "a sufficiently long backtest period and a beautiful return curve" as signals of a strategy’s effectiveness. Why do you think that's not enough?
Bernard zkBernard: Traditional intuition suggests that the longer the backtest, the better: if a strategy has been effective over the past three, five, or ten years, it may continue to be effective in the future.
However, when strategy design, explanation, or validation itself relies on large models, the situation changes. Models have a knowledge cutoff date and have encountered a plethora of publicly available market information, historical events, and existing narratives during training. If we validate a model's judgments using historical data prior to its knowledge cutoff, we are likely allowing it to handle a "test it has already taken."
Its high score does not necessarily indicate that this logic can be transferred to the future; sometimes, it is merely recalling or restating its known market memory.
TECHUB NEWS: How does this "open-book exam" risk manifest itself?
Bernard zkBernard: There are at least three typical scenarios.
The first: the model designed a strategy that yields suspiciously beautiful backtest results. It may already know the reasons behind the rise of a certain asset class over the past few years, and thus constructs a seemingly reasonable logic based on known outcomes.
The second: when researchers ask the model a different question, the conclusions received are almost unchanged. We might think this proves that the strategy is robust, but it could just be the same memory restated in a different way.
The third, which is more insidious: when researchers select a backtest period, they unconsciously know what has happened in that historical timeframe. For example, seeing a round of chip market activity and then specifically choosing this period to validate related strategies. At this point, the "open-book exam" applies not only to the model but also to the researcher themselves.
Therefore, the question is not just "Will AI make errors?" but whether the research process itself has brought known answers into the validation phase.
Not judging strategy correctness, but first auditing evidence
TECHUB NEWS: What exactly does your "Evidence Deflator" aim to solve?
Bernard zkBernard: It is not a tool that directly tells you "this strategy is valid" or "this strategy is invalid"; it is more like adding an auditing layer before the strategic conclusion.
I hope it first checks several key issues: what is the knowledge cutoff date of the model itself; what is the start and end time of the backtest; what statistical indicators and thresholds have researchers set; and how many sets of parameters were actually tried to obtain the final result, and what parameter search structure was adopted.
This information collectively determines how much real, independent evidence a seemingly long historical sample can provide. The tool aims to convert "surface evidence" into "actual evidence" and give researchers a clear judgment and next steps.
TECHUB NEWS: In other words, the focus is not on whether the backtest is "good-looking," but whether it has sufficient resolution?
Bernard zkBernard: Yes. These two aspects are very different.
Assuming a backtest spans 66 months, but most of that history may have already been covered by the model’s training data, the truly uncovered, relatively independent sample may be only 8 months; and according to the risk and statistical characteristics of this strategy, it might require about 48 months to have adequate statistical resolution.
At this point, it is not satisfactory to simply say "the conclusion is not strong enough"; a more accurate statement is: this backtest currently does not have sufficient resolution. It has not provided enough independent samples for us to judge whether the strategy genuinely possesses transferable effectiveness.
This distinction is crucial. "The conclusion is not strong enough" easily leads people to continue seeking reasons for their original viewpoints; "the evidence lacks resolution" requires researchers to return to the validation design itself and acknowledge that they cannot conclude at this moment.
The more parameters adjusted, the longer the required validation
TECHUB NEWS: Besides model knowledge coverage, why are you particularly concerned about repeated parameter tuning?
Bernard zkBernard: Because it is quite common to explore parameter adjustments, but if researchers try many sets of parameters and ultimately only present the best-performing set, it will result in significant selection bias.
For example, a strategy might originally require about 48 months of effective validation period; if 20 sets of parameters were tried during the process and only the best results were retained, the needed validation period could be significantly raised to about 112 months. The reason is simple: the "optimal result" may partly arise from randomness rather than the actual capability of the strategy.
The most important reminder for ordinary researchers is: do not only retain the final winning parameter combinations. The number of parameter trials, screening rules, and eliminated schemes are also part of the research evidence and should be recorded.
TECHUB NEWS: If the tool indicates "insufficient evidence" or "backtest lacks resolution," what should researchers do next?
Bernard zkBernard: It is not about mechanically extending the backtest duration but taking action based on where the problem lies.
If there is not enough independent sample, stricter out-of-sample testing should be maintained, or one should wait for and gather more new data that has not been covered by existing knowledge; if there are too many parameter searches, one should reduce unrestricted trial-and-error and comprehensively record the parameter testing process; if the historical data itself is insufficient to support the judgment, one can also consider a different validation approach, such as cross-market, cross-asset class, or testing under different market conditions.
The value of the tool lies not in replacing researchers’ judgments but in making "currently untrustworthy" an explicable and actionable conclusion, rather than a vague risk warning.
Returning from "how to use" to "why it is credible"
TECHUB NEWS: Why did you participate in the AI research investment elite incubation program with this project?
Bernard zkBernard: On one hand, I want to meet more peers working in AI research investment to see how they understand research, strategies, and products; on the other hand, I want to validate the idea: is there truly a need for a more proactive auditing layer in the market?
Traditional funds engage third parties to verify performance, such as establishing verifiable records around investment performance standards. But I’m wondering, why must we wait until performance is out, or even when problems arise, to do verification? Can we not leave a verifiable record of the verification process before the strategy and research methods reach a conclusion?
I aim to push this auditing logic one step forward.
TECHUB NEWS: During the incubation process, did you receive any feedback that made you rethink the product?
Bernard zkBernard: Yes. Mentors and peers continuously ask: what exactly is this mechanism? Why can these numbers help identify whether the backtest is valid?
Upon reflection, product developers can easily explain "how to use the product" as "why the product is valid": what I input and what the system outputs, and how to measure it. But what users really want to know is why these parameters can identify model recall, sample coverage, or selection bias.
This reminds me that a product must not only have processes but also clearly articulate its principles. Especially in AI research investment, if a tool itself cannot be understood or revisited, it is difficult to become a truly trustworthy infrastructure.
Future: Embed in workflow, not replace judgment
Currently, this tool is being trialed in a small scope among colleagues and friends, mainly to test the logical validity of individual investment strategies. Bernard also maintains a cautious attitude: even if individual positions have performed well recently, it may not entirely stem from the strategy itself; it could merely coincide with an overall market rise.
"If the strategy only follows the market, then I could just buy the index," he said. The next step is to compare the strategy with the S&P 500, QQQ, and other indices closer to the sector of the strategy to further differentiate the potential Alpha generated by the strategy from the returns brought by the overall market Beta.
In product iteration, he hopes to prioritize strengthening the evaluation of sample independence, such as measuring the overlap between different parameter variants; meanwhile, he is also considering integrating the tool into backtesting and agent workflows via API or MCP, reducing dependence on a standalone front-end interface. The initial target users will primarily be institutions with risk control and strategy optimization needs, while also considering individual researchers.
For those just beginning to use large models for investment research, strategy generation, or backtesting, Bernard hopes they first ask themselves one question:
"How much of this beautiful result comes from truly independent, sufficient, and credible evidence?"
As AI makes "creation" cheaper, validation may become a more scarce ability. The "Evidence Deflator" does not provide buy or sell conclusions; rather, it seeks to offer a more cautious path: first confirming whether a conclusion can withstand validation before discussing whether the strategy is credible.
Editor's Note: This article is based on TECHUB NEWS's recordings, meeting minutes, and interview outlines with Bernard zkBernard. The company, product functions, technical solutions, and subsequent plans mentioned are based on his interview statements and interview materials; the sample periods, number of parameters, and validation periods involved in tool examples are used solely for illustrative purposes and do not represent actual investment performance or future results.
Disclaimer: This article is for information exchange and discussion of research methods only and does not constitute any investment advice. Prices of securities, futures contracts, and virtual assets may rise or fall, and past performance does not guarantee future results. Readers should not rely solely on the contents of this article to make investment decisions and should prudently assess their own investment goals, financial situation, and risk tolerance, consulting independent professional advice when necessary.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。
