NanoML Lab 私密原型 / Private prototype
INTERACTIVE LEARNING PROTOTYPE · 交互式教学原型

从分子特征到模型输出, / From molecular features to model outputs,
看见每一步发生了什么 / see what happens at each step

调整分子描述符、Morgan 指纹设置与 Residual MLP 结构,观察输入表示如何穿过网络并改变教学预测结果。 / Adjust molecular descriptors, Morgan fingerprint settings, and the Residual MLP architecture to see how input representations pass through the network and change the demonstration output.

进入端到端研究工作台 → / Open the end-to-end research workbench → 开始探索 / Start exploring 无需上传数据 · 特征探索本地计算 / No data upload required · Feature exploration runs locally
C
O
N
C
01特征编码 / Feature encoding
02残差学习 / Residual learning
03教学输出 / Demonstration output
教学演示 / Educational demonstration 本原型不使用实验训练数据,输出仅用于理解模型机制,不代表真实药物、纳米组装或实验成功概率。 / This prototype uses no experimental training data. Outputs illustrate model mechanisms and do not represent real drug performance, nanoparticle assembly performance, or experimental success probabilities.
01 / MODEL EXPLORER

端到端模型实验台 / End-to-end model sandbox

1
选择分子 / Select a molecule教学预设 / Tutorial presets
2
构建表示 / Build representations指纹 + 描述符 / Fingerprints + descriptors
3
配置网络 / Configure the network残差多层感知机 / Residual MLP
4
观察输出 / Inspect outputs教学分数 / Demonstration score
FEATURE REPRESENTATION

模型“看到”的分子 / The molecule as the model “sees” it

Morgan 指纹 · 64-bit 教学视图 / Morgan fingerprint · 64-bit tutorial view15 个 ON bits / 15 active bits
检测到片段 / Fragment detected未激活 / Inactive

这里用可重复的哈希规则压缩展示“局部原子环境 → bit”的思想;不是 RDKit 计算结果。 / Reproducible hashing rules provide a compact illustration of “local atomic environment → bit”; these are not RDKit-computed results.

标准化描述符向量 / Standardized descriptor vector5 维 / 5 dimensions

标准化把不同单位放到相近尺度,避免数值大的特征仅因单位而主导模型。 / Standardization brings features with different units onto similar scales, preventing features from dominating merely because their numerical values are larger.

当前进入 MLP / Current MLP input 64-bit 指纹 ⊕ 5 个描述符 → 69 维融合向量 / 64-bit fingerprint ⊕ 5 descriptors → 69-dimensional fused vector
RESIDUAL MLP

信息如何穿过网络 / How information passes through the network

残差已启用 / Residual connections enabled
为什么加残差? / Why add residual connections? 主路径学习新的特征组合;捷径把较早的信息直接送到后面。两条路径相加,深层网络更容易保留有用信号。 / The main path learns new feature combinations, while the shortcut carries earlier information directly forward. Adding the two paths helps deeper networks preserve useful signals.
TEACHING OUTPUT

教学预测与关联解释 / Demonstration predictions & explanations

非实验预测 / Not an experimental prediction
0.68教学分数 / Demonstration score
中等偏高的演示输出 / Moderately high demonstration output

当前组合在教学规则中较为平衡 / The current combination is relatively balanced under the tutorial rules

这只是用于观察输入—模型—输出关系的合成函数,不能用于筛选真实实验候选。 / This synthetic function only illustrates the input–model–output relationship and cannot select real experimental candidates.

以姜黄素预设为起点 / Starting from the curcumin preset
哪些输入正在推动这个演示值? / Which inputs influence this demonstration value?方向与强度 / Direction & strength
PUBMED × WHO ATC × ChEMBL

从疾病出发,找到可核查的药物线索 / Find verifiable drug leads starting from a disease

研究线索 · 非治疗建议 / Research leads · Not treatment advice

PubMed 提供文献,ATC 提供药物分类。名称出现在摘要中,不代表研究结果支持该药物,更不代表它能够自组装成纳米颗粒。 / PubMed provides literature; ATC provides drug classifications. A drug name appearing in an abstract does not establish support for that drug, nor demonstrate its ability to self-assemble into nanoparticles.

支持常见癌症中文名称;其他疾病请用英文。按 PubMed 相关性返回前 20 篇,不是系统综述。无需 智谱 密钥即可检索。 / Common cancer names in Chinese are supported; use English for other diseases. The top 20 papers are returned by PubMed relevance. This is not a systematic review. No Zhipu API key is required for searching.

等待检索。查询词会发送给 NCBI PubMed。 / Awaiting a search. Search terms will be sent to NCBI PubMed.

ChEMBL · 核查分子结构与性质 / ChEMBL · Verify molecular structures & properties

输入英文药名或 ChEMBL ID。名称搜索最多显示 5 个匹配条目,请核对游离体、盐型和结构。缺失数值保留为空,不以 0 代替。 / Enter an English drug name or ChEMBL ID. Name searches show up to 5 matches. Verify the free form, salt form, and structure. Missing values remain missing rather than being replaced with zero.

公开数据查询无需 智谱 密钥。 / Public-data queries do not require a Zhipu API key.

本区读取结构和描述符;端到端工作台可查看实验活性、靶点与作用机制记录。ChEMBL 的 ATC 字段为交叉引用,分类现状请核对 WHO 官方索引。此处不将数据库条目自动转换为纳米颗粒实验预测。 / This section retrieves structures and descriptors. The end-to-end workbench shows experimental bioactivity, targets, and mechanisms of action. ChEMBL ATC fields are cross-references; verify current classifications in the official WHO index. Database records are not automatically converted into nanoparticle experimental predictions.

ATC 在这里能做什么?覆盖哪些药物? / What does ATC do here, and which drugs are covered?

ATC 根据解剖、治疗与化学特征组织药物分类,不提供分子指纹,也不证明某一疾病的疗效。本版人工核对多柔比星 L01DB01、紫杉醇 L01CD01、顺铂 L01XA01、氟尿嘧啶 L01BC02,核对日期 2026-09-15。只匹配返回标题/摘要中的英文通用名及有限别名,不做全库实体识别;缩写与同名仍需人工核查。 / ATC classifies drugs by anatomical, therapeutic, and chemical characteristics; it neither provides molecular fingerprints nor establishes efficacy for a disease. This version manually checks doxorubicin L01DB01, paclitaxel L01CD01, cisplatin L01XA01, and fluorouracil L01BC02, verified on 2026-09-15. Matching is limited to English generic names and selected aliases in returned titles and abstracts, not database-wide entity recognition. Abbreviations and ambiguous names still require manual review.

打开 WHO 合作中心 ATC/DDD 索引 ↗ / Open the WHO Collaborating Centre ATC/DDD Index ↗
LLM + PDF + RAG

让模型带着“证据”回答 / Let the model answer with evidence

PDF 解析负责把文件变成文字,RAG 负责找到相关片段,LLM 最后依据片段组织答案。三者不是同一个功能。 / PDF parsing turns a file into text, RAG finds relevant passages, and the LLM uses those passages to compose an answer. These are distinct functions.

智谱 接口已预留 / Zhipu integration prepared未配置密钥时仅展示检索原文 / Only retrieved source text is shown without an API key
PDF文档输入 / Document input上传或示例 / Upload or example
→
01解析文本 / Parse textPDF.js / 可替换 MinerU / PDF.js / replaceable with MinerU
→
02切分与索引 / Chunking & indexing重叠文本块 / Overlapping text chunks
→
03检索证据 / Retrieve evidence相关度排序 / Rank by relevance
→
04LLM 生成 / LLM generation基于上下文 / Based on context
✓解析 / Parse 2切块 / Chunk 3检索 / Retrieve 4生成 / Generate
AI
RAG 助手 / RAG assistant等待提问 / Awaiting a question

选择一个示例问题,或上传带有文字层的 PDF。回答会同时展示检索到的来源片段,方便你核对依据。 / Select an example question or upload a PDF with a text layer. Retrieved source passages are shown alongside the answer so you can check the evidence.

原型边界 / Prototype limitations PDF 在本地解析;点击问答会将问题和检索片段提交至站点服务端,启用 LLM 后发送给 智谱;本地词汇检索可直接运行(非语义向量检索)。真实 LLM 生成接口已按服务端密钥方式预留,缺少密钥时不会伪装成模型回答。 / PDFs are parsed locally. Submitting a question sends the question and retrieved passages to the site server and, when enabled, to Zhipu. Local lexical retrieval works directly and is not semantic vector retrieval. LLM generation is prepared for a server-side API key; without a key, source excerpts are not presented as model-generated answers.
02 / CONCEPT MAP

把三个概念连成一条线 / Connect the three concepts

先认清“输入是什么”,再理解“模型做了什么”,最后谨慎解释“输出意味着什么”。 / First identify the inputs, then understand what the model does, and finally interpret what the outputs mean with care.

101 01 · STRUCTURE

Morgan 指纹 / Morgan fingerprints

像给分子结构做“片段打卡”。模型看到的是哪些局部结构出现过,而不是一张分子图片。 / Think of this as checking off fragments in a molecular structure. The model sees which local structures occur, rather than a picture of the molecule.

什么时候有用? / When is this useful?

当化学结构相似性、官能团或局部原子环境可能影响任务时,它提供紧凑的结构信号。 / It provides compact structural information when chemical similarity, functional groups, or local atomic environments may matter for the task.

x̄ 02 · PROPERTY

分子描述符 / Molecular descriptors

把分子量、亲脂性、极性等性质变成数值。它们比指纹更容易直接解释。 / Descriptors express properties such as molecular weight, lipophilicity, and polarity numerically. They are easier to interpret directly than fingerprints.

为什么要标准化? / Why standardize?

MW 可能是几百,HBD 只有个位数。标准化后,模型不会只因为数值尺度大就偏爱某个特征。 / MW may be in the hundreds while HBD is a single-digit count. Standardization prevents a feature from being favored merely because its numerical scale is larger.

↷ 03 · LEARNING

残差多层感知机 / Residual MLP

MLP 学习特征的非线性组合;残差捷径让原始信息可以绕过部分变换再相加。 / An MLP learns nonlinear combinations of features. Residual shortcuts let earlier information bypass some transformations before being added back.

它不是 ResNet 图像模型吗? / Is this the same as a ResNet image model?

“残差”是一种连接思想,并非图像专属。把它放进全连接网络,就得到常说的 Residual MLP。 / Residual connections are an architectural idea, not exclusive to images. Adding them to a fully connected network produces what is commonly called a Residual MLP.

!
从原型走向真实研究,还缺什么? / What is still needed to move from prototype to research?

需要经过清洗且与目标任务匹配的实验数据、合理的数据划分、基线模型、外部验证、不确定性评估,以及最终的湿实验验证。 / Clean experimental data matched to the target task, appropriate data splits, baseline models, external validation, uncertainty assessment, and ultimately wet-lab validation are needed.