マテリアルズインフォマティクス・材料
AIによる材料発見——ふるいには効くが発見者にはまだ早い
目次
機械学習は、材料探しの「安いふるい」として候補を絞るのに効く。 では、その実力はどこまでか。近年の華々しい見出しを、冷静に検証してみよう。
ここ数年、「AIが材料を発見した」というニュースが立て続けに出た。そのたびに、専門家が中身を見にいき——多くが「実は新しくなかった」ことが分かってきた。これは、AIが材料探しで効く場所と盛りすぎる場所を、はっきり教えてくれる。
まず、言葉を分ける——「安定」≠「作れる」≠「役立つ」≠「新しい」
つまずきの根っこは、4つの別々のことを、つい一緒くたにしてしまうことだ。
「計算で安定そう(①)」は、ただの出発点だ。そこから「実際に作れる(②)」「役立つ性質を持つ(③)」「既知でなく本当に新しい(④)」までには、何段もの関門がある。この関門の連なりは、候補を安いふるいから順に絞り込む漏斗——安いふるいを先に、高いふるいを後に——そのものだ。 ところが見出しは、①の数を「新材料の数」として語ってしまう。①の数を「新材料の数」として受け取る読まれ方が、近年の3つの出来事すべての火種になった。
3つの実例——華々しい発表と、その後の検証
① グーグルDeepMindの「GNoME」(2023年) 深層学習で、それまでの基準(Materials Project)に照らしてすでに凸包より下だと判定された構造を220万個見つけたと発表。新しい発見どうしも安定を競うので凸包は引き直され、その更新後の凸包の上に残る新規材料は約38万である。引き直しは過去にも起きていて、GNoME 自身が Materials Project と OQMD の「安定」材料を少なくとも5000件、凸包の下へ追い出している。計算で安定とされていた材料は約4万8千から42万1千へ、ほぼ一桁拡大したという1。論文は新規性も数で主張している。発見した安定結晶は、既知のプロトタイプの置換や列挙からは出てこない4万5500超の新規プロトタイプに対応し、Materials Project の8000に対して5.6倍だという1。 ——だが2024年、専門家(Cheetham & Seshadri)がChemistry of Materials誌で反論した。彼らの結論は「新規性・確からしさ・有用性を同時に満たす化合物の証拠は乏しい」で、繰り返し挙げられるのは二点だ。ひとつは既知構造との実質的な同一性——同じ構造型を対称性だけ下げて別のものとして登録している例。もうひとつは、実際には乱れているはずの金属イオンを整然と並べた構造がデータベース全体を通じて繰り返し現れることだ(0Kの密度汎関数計算はエントロピーを数えないので、秩序のある解を選んでしまう)。この偏りには規模も示されている——予測側では上位4つの空間群がすべて非中心対称で、全体の約34%を占める。ICSD では上位24位までに非中心対称の空間群は1つしかなく、1%にとどまる。彼らが38万4870件から無作為に抜き出した10件は、10件とも ICSD に既知の構造として見つかった。そのうえで、どれにも機能が示されていない以上「材料(material)」とは呼べず、正しくは結晶性無機化合物の提案リストだ、というのが彼らの立場である——これは冒頭に挙げた関門③そのものだ。実際この論考は、放射性のプロメチウム・アクチニウム・プロトアクチニウムを含む化合物が1万8138件、テクネチウム・ネプツニウム・プルトニウムを含むものがさらに2万3529件あり、これらは材料として使える見込みが無いと書いている。ただし同じ論考は、手法そのものは健全だとも書く。新しい組成の多くは既知材料の些細な焼き直しだが、計算手法は全体として妥当な組成を出しており、その基盤は健全だという確信を与える、と。彼らはさらに、予測を既知の文献と突き合わせて本当に新しくないものを濾す作業が要るとし、40万件近い一覧に対してそれを人手で行うのは非現実的だと書いている2。
② バークレーの「A-Lab」(2023年) AIとロボットの自律実験室が、17日で58の標的のうち41の新化合物を合成した、とNatureに発表3。 ——批判は2024年3月に出た。プリンストンとUCLの7人(筆頭はジョシュ・リーマン、責任著者にレスリー・ショープとロバート・パルグレイブ)が、X線データの解析が甘く、多くは既知の物質の取り違えだと指摘する。彼らの結論は「この研究で新材料はひとつも発見されていない」だった。正しく合成されたと認められるのは3件だけで、それなら成功率は58標的中の3件——5%になる。その内訳として、約3分の2は予測された規則構造ではなく既知の乱れた相だった可能性が高い、というのが彼らの読みだ4。 最終的にNatureは訂正(Author Correction)を出した(2026年1月オンライン)。論文タイトルから「新規(novel)」の語を削り、訓練データに紛れていた1化合物(Zn₂Cr₃FeO₈)を除外した——これで標的は58→57、報告された成功も41→40件になる。そのうえで訂正後の記載は「57標的のうち36化合物」である。40件のうち36件は判定が正しく、残る4件はX線回折だけでは判定不能——これが著者らの手作業による再解析の結論で、この再解析そのものが公開後にあらためて査読を受けている(Nature は訂正の謝辞で、再解析後の回折データを評価した査読者に触れている)。ここで言う「新規」も「科学的に新しい」ではなく「予測プラットフォームにとって新しい」の意味だった、と著者自身が認めている。同じ弁明は2024年の批判論文が著者らの応答としてすでに記録しており、批判側は、それでも「新規材料」を掲げる題名と要旨は大いに誤解を招くと切り返していた4。撤回ではなく訂正。ただし批判側の数え方(58標的中3件)と訂正後の数え方(57標的中36件)は、今も揃っていない。揃わないのは、成功の定義が違うからだ。A-Lab 論文は、目標と同じ結晶構造と組成を保つなら部分的に乱れた版ができても成功と数えると Methods に書いており、訂正後の版ではこの定義を本文の冒頭にも置いた3。批判側は、既知の乱れた相ができたなら予測した規則構造は作れていないと数える。同じ回折データが、片方では成功に、片方では失敗に読まれている。新規性の基準も同じ形で割れる。批判側は、A-Lab が新規性の基準を明示していたことを記録している。標的を Materials Project で「理論上のみ」(ICSD に無い)と印の付いた化合物から選ぶ、という基準だ。そのうえで批判側は、この基準には異論がありうると断る。Materials Project には組成の乱れた化合物が無く、既知でも ICSD に無い化合物、とくに乱れたものは多いからだ。それでも自分たちの評価にはおおむねこの基準を使う、と彼らは書いている4。リーマンらが訂正後に応答した記録は、ここで引いた2出典には無い4。
③ マイクロソフトの「MatterGen」(2025年) 既存候補を“ふるう”のではなく、新しい結晶を生成する拡散モデル。実証として設計・合成された化合物 TaCr₂O₆ が話題になった。ただしMatterGen論文自身が、実際に作れたのは4候補のうち1つで、しかも出てきたのは予測した規則構造そのものではなくその組成が乱れた版だと書いている5。 ——だが2026年、Mikkel Juelsholt がMaterials Horizons誌(査読付き)で指摘した。合成された乱れた相 Ta₁/₃Cr₂/₃O₂ は、1971年に報告済みの Ta₁/₂Cr₁/₂O₂ と——c軸まわりに90°回せば——同一構造である。Juelsholt 自身、Ta₁/₃Cr₂/₃O₂ という組成そのものは報告例が無いと認めたうえで、それは Ta の含量が0.65以下ならルチル構造が CrO₂ まで続く既知の固溶体系列の一点にすぎず、真の新化合物とは言えないと退けている。しかもその ICSD 登録(collection code 9516)は、MatterGen が新規性を判定するのに使う参照データセットに entry 956681 として入っていた。訓練データの側にも、同じ Ta–Cr 酸化物のルチル構造をP1へ展開したものが複数入っている(Materials Project の mp-753467 と mp-756340、Alexandria の agm003743920。訓練データ上の表記は TaCrO4)。つまり「新発見」ではなく、手元のデータの言い換えだ、というのだ。 Juelsholt は性質の側でも切っている。この実証はもともと体積弾性率200 GPa を狙って生成したものだが、Zeni らの値は、ナノインデンテーションで測った Young 率から推定した体積弾性率で、4回の測定の最大値が169 GPa である(平均は158 ± 11 GPa。粉末試料が緻密でない可能性から、最大値を最良の推定としている)5。既知の Ta₁/₂Cr₁/₂O₂ は焼結体で181 GPa と報告されており、Juelsholt はそちらのほうが硬いと切る。「では二番目に硬い未知の組成を当てたのでは」という擁護も、彼は二点で否定する。Ta₁/₂Cr₁/₂O₂ に近い組成なら硬さもほぼ同じはずでそちらを予測すべきだったこと、そして訓練データに入っていたのは化合物そのものだけで硬さの値ではないことだ。なお、この指摘は公表前に Zeni らへ連絡しないまま出されたと、著者自身が断っている6。
公平に——では、AIは役立たずなのか?
いや、それは違う。
機械学習は、既知の化学の“内側”で、性質を高速に見積もって候補を絞ること(=スクリーニングの“安いふるい”の段)には、実際に強い。GNoME 論文はこの段の性能を的中率で書いている。予測した候補のうち第一原理計算で安定と確かめられた割合は、構造を与えた場合で80%超、組成だけの場合で33%で、先行研究の1%から上がった1。ただしこれは計算上の安定(①)に対する的中率であって、作れるか(②)の的中率ではない。GNoMEの予測のうち736件は実験構造データベース(ICSD)の登録構造と一致したと報告されている。ただし論文自身が Methods で、そのうちプロジェクト開始後に新たに報告された物質に対応するのは184件だと書いている——残りは以前から知られていた物質との一致だ1。それでもスクリーニングの加速そのものは本物だが、この736という数字を「発見」の証拠として読むことはできない。
問題は、その強みが「内挿(interpolation)」——学んだデータの周辺をなめらかに埋めること——に偏っていることだ。 一方で見出しが期待させるのは「外挿(extrapolation)」——誰も見たことのない、本当に新しい化学へ大きく踏み出すことだ。ここでのAIの弱さは「まったくできない」という単純な話ではない。一般には、訓練データから遠ざかるほど予測の確からしさは落ちるとされる。だから“新発見”ほど当たり外れが大きくなる。ただしGNoME論文自身は、規模を上げると訓練で見ていない元素数の領域へ一般化する能力が現れる(emergent out-of-distribution generalization)と主張しており、ここは決着していない1。上で見たGNoMEやA-Labの“発見”が検証の段で次々に既知物質へと差し戻された主因として批判側が挙げるのは、内挿か外挿かではない。性質の似た元素が同じ結晶サイトを共有する「乱れ」を予測が扱えず、既知の乱れた相を新しい規則構造と取り違えたこと、そしてX線回折の自動解析が同定に耐えなかったことだ(内挿/外挿という整理は本稿のもので、批判側の診断ではない)。乱れの取り違えは GNoME 批判の側も名指ししていて、彼らは同じ問題が A-Lab を論じた別の論考でも指摘されていると書き添えている2。外挿の弱さと合わせて、“新発見”を名乗るにはなお検証が要るということだ。
批判の側は、その処方まで書いている。Juelsholt は、MatterGen が自分の訓練データと自分の予測を区別できない以上、新規性を担保するには予測1件ごとに人が確かめるしかなく、それが売りだったはずの高速な予測を妨げる、と述べる。そのうえで彼は、これを1つのモデル固有の欠陥ではなく、結晶学と原子の乱れという概念に生成AIツールが総じてつまずいている例として位置づけた6。
要するに——AIは、材料探しの速い“ふるい”としては本物だ。だが「発見者」と呼ぶには早い。「計算で安定」と「作って役立つ新材料」の間には、まだ人間の合成と検証が要る、深い谷がある。
出典6件
-
GNoME(Google DeepMind), A. Merchant, E. D. Cubuk ら「Scaling deep learning for materials discovery」, Nature(2023年11月)。先行研究(Materials Project)を基準にすれば安定な構造が220万個、うち38万1000件が更新後の凸包の上に残る新規材料——原典の逐語は
Of these, 381,000 entries live on the updated convex hull as newly discovered materialsで、38万1000は220万の部分集合である。ICSD の実験構造と一致したのは736件で、原典は要旨でこれを736 have already been independently experimentally realizedと書く。ただしこの「独立に」はGNoMEと無関係な他者の実験でという意味であって、予測が実験に先行したという意味ではない——著者らの Methods は184 of these structures correspond to novel discoveries since the start of the projectと書いており、予測が実験に先行したと読めるのはこの184件だ。原典はこの38万1000についても、将来の発見で凸包から落ちうると自ら書いている——Consistent with other literature on structure prediction, the GNoME materials could be bumped off the convex hull by future discoveries, similar to how GNoME displaces at least 5,000 'stable' materials from the Materials Project and the OQMD。新規性の側の主張はleading to more than 45,500 novel prototypes in Fig. 2c (a 5.6 times increase from 8,000 of the Materials Project), which could not have arisen from full substitutions or prototype enumeration(プロトタイプの計数は XtalFinder による)。ふるいの性能は的中率でimprove the precision of stable predictions (hit rate) to above 80% with structure and 33% per 100 trials with composition only, compared with 1% in previous workと書かれ、能動学習の出発点はstart at less than 6% and 3%だった。この的中率は DFT に対する安定判定の精度であって、合成可能性ではない。 https://www.nature.com/articles/s41586-023-06735-9 ↩ ↩2 ↩3 ↩4 ↩5 -
GNoME批判:A. K. Cheetham & R. Seshadri「Artificial Intelligence Driving Materials Discovery? Perspective on the Article: Scaling Deep Learning for Materials Discovery」, Chemistry of Materials(ACS, 2024)。「新規性・確からしさ・有用性という三拍子を満たす化合物の証拠は乏しい」(逐語:
scant evidence for compounds that fulfill the trifecta of novelty, credibility, and utility)と結論。繰り返し挙げるのは、既知構造と実質同一で対称性だけが低い登録(These structures are virtually identical, aside from the lower symmetry in the GNoME entry/pseudosymmetry that is present in many of the entries)と、many of the entries are based upon the ordering of metal ions that are unlikely to be ordered in the real worldの二点。無作為抽出(selecting 10 compounds from among the 384,870 database entries)についてはWe were able to identify the structure of every one of the 10 Stable Structure entries in the ICSD database, albeit usually with a space group of higher symmetry than that in the AI database、「材料」という語についてはSince no functionality has been demonstrated for the 384,870 compositions in the Stable Structure database, they cannot yet be regarded as materials。空間群の偏りは数でも示されていて、In fact, the top four space groups in the Stable Structure database are all non-centrosymmetric and account for ∼34% of all the structures. By contrast, in the ICSD there is only one non-centrosymmetric space group in the top 24 and it accounts for only 1% of all structures。乱れの見落としが A-Lab 批判と共通することも著者ら自身が書いている(The main reason for the disparity is due to the frequent prediction of structures with atoms ordered on distinct crystallographic sites that—in most cases—are likely to be disordered/We also note that this issue has been highlighted+in a commentary on a different Nature article on robotic/AI-based materials discovery。参照先は Leeman ら PRX Energy 3, 011002 と Szymanski ら Nature 624, 86–91)。放射性元素についてはIn fact, there are 18,138 compounds of such radioactive elements in the large Stable Structure database, including those of Pm, Ac, and Pa/There are a further 23,529 entries for compounds containing the highly radioactive elements Tc, Np, and Pu、結論節はThese include the elimination of a large number of radioactive materials that are unlikely to have any utility in the materials worldと書く。濾し分けの負担についてはWhat is now needed is greater effort to connect the predictions to what is already known in the literature to filter out the many candidates that are not truly novel. It is impractical to do this manually with a list of almost 400,000 new compositions。一方でthe computational approach delivers credible overall compositions, which gives us confidence that the underlying approach is soundとも書いており、手法そのものを否定してはいない。 https://pubs.acs.org/doi/10.1021/acs.chemmater.4c00643 ↩ ↩2 -
A-Lab(バークレー):N. J. Szymanski ら「An autonomous laboratory for the accelerated synthesis of novel materials」, Nature 624, 86–91(2023年11月)。当初「17日・58標的のうち41新化合物」。この表題は原版のもので、2026年の訂正で “novel” が “inorganic” に改められた。このURLが現在返すのは訂正後の版であり、表題は
An autonomous laboratory for the accelerated synthesis of inorganic materials、要旨はrealized 36 compounds from a set of 57 targetsになっている。成功の定義は原版の Methods にSuch cases were still considered successful as long as the disordered version of the target retained the same crystal structure and overall composition as the ordered versionと在り、訂正後の版はこれを本文の冒頭節にもwe consider a synthesis procedure successful when it forms either an ordered or partially disordered version of its target materialとして置いた(原版の本文には無い文)。 https://www.nature.com/articles/s41586-023-06734-w ↩ ↩2 -
A-Lab批判と訂正:(批判) J. Leeman, Y. Liu, J. Stiles, S. B. Lee, P. Bhatt, L. M. Schoop, R. G. Palgrave「Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis」, PRX Energy 3, 011002(2024)。結論は
These errors unfortunately lead to the conclusion that no new materials have been discovered in that work(要旨)/we believe that at time of publication, none of the materials produced by A-lab were new(序論)。成功率についてはwe could agree that three materials were correctly synthesized … In this case, the success rate would be 3/58, or 5%、内訳はtwo thirds of the claimed successful materials in Szymanski et al. are likely to be known compositionally disordered versions of the predicted ordered compounds。A-Lab の新規性の基準については §II でIt seems as if the criterion for novelty of a material is its presence in the Materials Project (MP) and its absence from the ICSD. This criterion is open to criticismと述べたうえでwe will mostly use this criterion to assess the novelty of the A-lab synthesis productsと断り、§VI の勧告でもThis was clearly done by the A-lab paper (absence from the ICSD) but some may take issue with this definitionと書く。著者らの応答についてはEven if those known materials were not in the training set and thus were "new to the A-lab," as claimed by the authors in response, this makes the tile and abstract, which claim "novel materials," highly misleading(tileは原典の誤植)。 https://link.aps.org/doi/10.1103/PRXEnergy.3.011002 / (訂正) Nature の Author Correction(タイトルから novel を削除、2026年2月号 650:E1、s41586-025-09992-y):訓練データ混入の Zn₂Cr₃FeO₈ を除外、報告40成功のうち36が正しく4件はXRDのみでは判定不能、「新規」は予測プラットフォームにとっての新規の意。撤回ではなく訂正。この再解析は著者らが手作業で行ったものだが、訂正はThis re-analysis was peer-reviewed post-publicationと明記し、謝辞でもthe reviewers who assessed the re-analyzed PXRD data post-publicationに触れている。 https://www.nature.com/articles/s41586-025-09992-y ↩ ↩2 ↩3 ↩4 -
MatterGen(Microsoft Research):「A generative model for inorganic materials design」, Nature(2025年1月)。Materials Project と Alexandria から再計算した607,683の安定構造(20原子以下)で訓練した拡散モデル(
607,683 stable structures with up to 20 atoms recomputed from the Materials Project (MP) and Alexandria datasets)。訓練データは実験で確かめられた物質そのものではなく、DFT で再計算された構造である(Alexandria は計算生成のデータベース)。乱れた実在相が、計算の都合で規則構造として入っていることがある。新規性の判定に使う参照データセットはこれとは別で、ICSD を含むAlex-MP-ICSD, comprising 850,384 unique structuresのほうである。実証の4件は無作為に選ばれたのではない——原典は候補を(1) uniqueness and novelty; (2) energy above the hull stability from MatterSim and DFT; (3) phonon stability from MatterSim; and (4) whether the material contains oxygenで絞って75件にし、from which we select four for experimental synthesis after expert inspectionと書いている。つまりこの4件は、新規性の篩と専門家の目視を通り抜けたものだった。合成が成功したのはそのうち1件で、Synthesis was successful for one of the four candidates。できたものを著者らはthe synthesized material is TaCr2O6, a compositionally disordered version of the ordered structure predicted by MatterGenと記述している。体積弾性率はWe also experimentally measure the Young's modulus of the sample by nanoindentation and estimate its bulk modulus using the DFT-computed Poisson ratio of 0.30. The estimated bulk modulus is up to 169 GPa after four measurements (158 ± 11 GPa), in which the maximum of the four measurements is our best estimate given that the experimental powder sample is likely non-compactで、規則構造への DFT 予測値は 222 GPa。 https://www.nature.com/articles/s41586-025-08628-5 ↩ ↩2 -
MatterGen批判:Mikkel Juelsholt(単著)「Continued challenges in high-throughput materials predictions: MatterGen predicts compounds from the training dataset」, Materials Horizons(RSC, 2026, 査読付き)。要旨は
MatterGen was used to predict the novel compound TaCr2O6, which was subsequently synthesised in a disordered form as Ta1/3Cr2/3O2. However, … this is not a novel compound but is identical to the previously reported Ta1/2Cr1/2O2, first described in 1971。同一性はupon a 90° rotation around the crystallographic c-axis, the two structures become identicalとして示され、組成の未報告はWhile the exact composition of Ta1/3Cr2/3O2 has never been reported, it is just a point in an already established solid-solution series of Ta Cr oxidesと自ら認めたうえで退けている。181 GPa の出所(ref.34)は S. Zhang ら, J. Adv. Ceram. 13, 373–387 (2024) で、熱間プレス焼結した CrTaO₄ の測定値である。訓練データ混入はICSD collection code 9516 is listed in the reference data set used by MatterGen as entry 956681と同定子つきで裏づけられている。著者の総括はMatterGen cannot distinguish between its training data and the compounds it predicts. Therefore, human inspection of each predicted compound is needed to ensure novelty, which hinders the rapid materials prediction offered by MatterGen、およびMatterGen joins other generative AI tools that struggle with key materials science concepts such as crystallography and atomic disorder。参照データセットとは別に訓練データ側も名指ししている(Furthermore, the training data used for MatterGen contains multiple other Ta Cr oxide structures that are rutile structures expanded to P1/Two of the structures are from the Materials project, mp-753467 in Fig. 2A and mp-756340 in Fig. 2B, while the last is from the Alexandria datasets, agm003743920/(listed as TaCrO4 in the training data))。性質の面の指摘はTa1/3Cr2/3O2 could be considered novel because of its properties as Ta1/3Cr2/3O2 was predicted to be a superhard material. Zeni et al. measured the bulk modulus to 169 GPa/However, Ta1/2Cr1/2O2 is an already known superhard material with a bulk modulus of 181 GPa and therefore harder than the compound measured by Zeni et al./However, this argument fails for two reasons/Secondly, the MatterGen training dataset did not include the hardness of Ta1/2Cr1/2O2, only the material itself/MatterGen, therefore, appears to fail to correctly account for the compositional effects of hardness in this particular class of materials。反論機会については著者自身がZeni et al., who was not contacted before the publication of this paperと断っている。 https://pubs.rsc.org/en/content/articlelanding/2026/mh/d6mh00268d ↩ ↩2
この記事はAIが執筆しています。内容には誤りが含まれる可能性があります。ご注意ください。