<?xml version='1.0' encoding='UTF-8'?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
  <id>https://ericspencer.us/ai4fm/</id>
  <title>AI4FM Updates</title>
  <updated>2026-09-15T17:10:04.972809+00:00</updated>
  <link href="https://ericspencer.us/ai4fm/"/>
  <link href="https://ericspencer.us/ai4fm/blog/atom.xml" rel="self"/>
  <generator uri="https://ablog.readthedocs.io/" version="0.11.12">ABlog</generator>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/tla-prover-accepted-2026/</id>
    <title>TLA-Prover Paper Accepted to ICSOFT 2026</title>
    <updated>2026-06-05T00:00:00+00:00</updated>
    <author>
      <name>Arslan Bisharat</name>
    </author>
    <content type="html">&lt;section id="tla-prover-paper-accepted-to-icsoft-2026"&gt;

&lt;p&gt;We are pleased to announce that our latest research paper, &lt;strong&gt;“TLA-Prover: Verifiable TLA+
Specification Synthesis via Preference-Optimized Low-Rank Adaptation”&lt;/strong&gt; (Paper #131), has been
accepted for publication at &lt;strong&gt;ICSOFT 2026&lt;/strong&gt; (International Conference on Software Technologies)
in Porto, Portugal.&lt;/p&gt;
&lt;p&gt;This marks our second paper accepted to ICSOFT 2026, alongside our systematic evaluation study
on LLM-based TLA+ generation. TLA-Prover introduces a 20-billion-parameter model that
achieves a &lt;strong&gt;3.5× improvement&lt;/strong&gt; in semantic correctness over public baselines by combining
supervised fine-tuning with repair-based reinforcement learning (GRPO).&lt;/p&gt;
&lt;p&gt;&lt;a class="reference external" href="../../papers/tla-prover/"&gt;Read the full paper details&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/tla-prover-accepted-2026/"/>
    <summary>We are pleased to announce that our latest research paper, “TLA-Prover: Verifiable TLA+
Specification Synthesis via Preference-Optimized Low-Rank Adaptation” (Paper #131), has been
accepted for publication at ICSOFT 2026 (International Conference on Software Technologies)
in Porto, Portugal.</summary>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="ICSOFT" label="ICSOFT"/>
    <category term="LLMs" label="LLMs"/>
    <category term="Publication" label="Publication"/>
    <category term="TLA+" label="TLA+"/>
    <published>2026-06-05T00:00:00+00:00</published>
  </entry>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/tla-prover-2026/</id>
    <title>New Paper: TLA-Prover — Fine-Tuning LLMs for Verifiable TLA+ Specification Synthesis</title>
    <updated>2026-05-16T00:00:00+00:00</updated>
    <author>
      <name>Arslan Bisharat</name>
    </author>
    <content type="html">&lt;section id="new-paper-tla-prover-fine-tuning-llms-for-verifiable-tla-specification-synthesis"&gt;

&lt;p&gt;Building on our &lt;a class="reference internal" href="../posts/llm-tla-evaluation-2025/"&gt;&lt;span class="doc"&gt;evaluation study&lt;/span&gt;&lt;/a&gt; that showed the best public LLMs
achieve only 8.6% semantic correctness on TLA+ specification generation, we present
&lt;strong&gt;TLA-Prover&lt;/strong&gt;: a 20-billion-parameter model trained specifically for TLA+ specification synthesis.&lt;/p&gt;
&lt;p&gt;TLA-Prover combines supervised fine-tuning (SFT) on verified examples with repair-based
group-relative policy optimization (GRPO), where the model learns to fix its own rejected
specifications. TLC model checking provides the reward signal directly — no learned reward
model is needed. A four-tier grading scheme (Bronze, Silver, Gold, Diamond) measures output
quality, with Diamond requiring that the model’s correctness property is meaningful enough
for TLC to detect deliberate violations.&lt;/p&gt;
&lt;p&gt;TLA-Prover achieves &lt;strong&gt;30% pass&amp;#64;1 at Gold and Diamond&lt;/strong&gt; on a held-out 30-problem benchmark —
roughly &lt;strong&gt;3.5× the 8.6% untuned baseline&lt;/strong&gt;. A DPO ablation trained from the same SFT
checkpoint reaches 20% at Diamond.&lt;/p&gt;
&lt;p&gt;&lt;a class="reference external" href="../../papers/tla-prover/"&gt;Read the full paper details&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a class="reference external" href="index.html"&gt;See other posts&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/tla-prover-2026/"/>
    <summary>Building on our evaluation study that showed the best public LLMs
achieve only 8.6% semantic correctness on TLA+ specification generation, we present
TLA-Prover: a 20-billion-parameter model trained specifically for TLA+ specification synthesis.</summary>
    <category term="DPO" label="DPO"/>
    <category term="Fine-Tuning" label="Fine-Tuning"/>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="GRPO" label="GRPO"/>
    <category term="LLMs" label="LLMs"/>
    <category term="TLA+" label="TLA+"/>
    <published>2026-05-16T00:00:00+00:00</published>
  </entry>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/gcasr-2026-posters/</id>
    <title>Presented Two Posters at GCASR 2026</title>
    <updated>2026-05-11T00:00:00+00:00</updated>
    <author>
      <name>Arslan Bisharat</name>
    </author>
    <content type="html">&lt;section id="presented-two-posters-at-gcasr-2026"&gt;

&lt;p&gt;On May 11, 2026 we presented two posters at the 13th Greater Chicago Area
Systems Research Workshop (GCASR 2026).&lt;/p&gt;
&lt;section id="a-structured-benchmarking-dataset-for-tla-specification-reasoning"&gt;
&lt;h2&gt;A Structured Benchmarking Dataset for TLA+ Specification Reasoning&lt;/h2&gt;
&lt;p&gt;Authors: Arslan Bisharat, Eric Spencer, Khushboo Bhadauria, Anisa Ramos, Brian Ortiz,
Mohammed Abuhamad, Konstantin Laüfer, TaiNing Wang, George K. Thiruvathukal&lt;/p&gt;
&lt;p&gt;This is a currently working paper.&lt;/p&gt;
&lt;/section&gt;
&lt;section id="large-language-models-llm-and-temporal-logic-of-actions-tla-how-effective-are-llms-for-verification-systems"&gt;
&lt;h2&gt;Large Language Models (LLM) and Temporal Logic of Actions (TLA): How Effective are LLMs for Verification Systems&lt;/h2&gt;
&lt;p&gt;Authors: Brian Ortiz, Arslan Bisharat, Eric Spencer, Khushboo Bhadauria, Anisa Ramos,
Mohammed Abuhamad, Konstantin Laufer, TaiNing Wang, George K. Thiruvathukal&lt;/p&gt;
&lt;link rel="stylesheet" href="../../_static/css/gallery.css" /&gt;
&lt;div class="ai4fm-gallery"&gt;
  &lt;img class="main-img" src="../../_static/images/Image.jpeg" alt="AI4FM team" /&gt;
  &lt;div class="thumbs"&gt;
    &lt;img src="../../_static/images/Image.jpeg" data-full="../../_static/images/Image.jpeg" alt="Team photo" /&gt;
    &lt;img src="../../_static/images/PXL_20260511_152740012.jpg" data-full="../../_static/images/PXL_20260511_152740012.jpg" alt="Close-up 1" /&gt;
    &lt;img src="../../_static/images/PXL_20260511_164300028.jpg" data-full="../../_static/images/PXL_20260511_164300028.jpg" alt="Close-up 2" /&gt;
  &lt;/div&gt;
  &lt;div class="controls"&gt;
    &lt;button class="g-prev"&gt;Previous&lt;/button&gt;
    &lt;button class="g-next"&gt;Next&lt;/button&gt;
  &lt;/div&gt;
&lt;/div&gt;
&lt;script src="../../_static/js/gallery.js"&gt;&lt;/script&gt;&lt;p&gt;&lt;a class="reference external" href="index.html"&gt;See other posts&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/gcasr-2026-posters/"/>
    <summary>On May 11, 2026 we presented two posters at the 13th Greater Chicago Area
Systems Research Workshop (GCASR 2026).</summary>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="GCASR" label="GCASR"/>
    <category term="LLMs" label="LLMs"/>
    <category term="TLA+" label="TLA+"/>
    <published>2026-05-11T00:00:00+00:00</published>
  </entry>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/llm-tla-accepted-2026/</id>
    <title>Our Paper on LLM-Based TLA+ Specification Generation Has Been Accepted</title>
    <updated>2026-04-30T00:00:00+00:00</updated>
    <author>
      <name>Arslan Bisharat</name>
    </author>
    <content type="html">&lt;section id="our-paper-on-llm-based-tla-specification-generation-has-been-accepted"&gt;

&lt;p&gt;We are excited to announce that our paper &lt;strong&gt;Can LLMs Write Correct TLA+ Specifications?
Evaluating Natural-Language-to-TLA+ Generation&lt;/strong&gt; has been accepted for publication at &lt;strong&gt;ICSOFT 2026&lt;/strong&gt; (International Conference on Software Technologies) in Porto, Portugal.&lt;/p&gt;
&lt;p&gt;The paper presents the first systematic evaluation of LLM-based TLA+ specification
synthesis from natural language, evaluating 30 LLMs across eight families on a curated
dataset of 205 TLA+ specifications. Key findings include that LLMs achieve up to 26.6%
syntactic correctness but only 8.6% semantic correctness, and that model size does not
predict quality on formal languages.&lt;/p&gt;
&lt;p&gt;&lt;a class="reference external" href="../../papers/llm-tla-evaluation/"&gt;Read the full paper details&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/llm-tla-accepted-2026/"/>
    <summary>We are excited to announce that our paper Can LLMs Write Correct TLA+ Specifications?
Evaluating Natural-Language-to-TLA+ Generation has been accepted for publication at ICSOFT 2026 (International Conference on Software Technologies) in Porto, Portugal.</summary>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="LLMs" label="LLMs"/>
    <category term="Publication" label="Publication"/>
    <category term="TLA+" label="TLA+"/>
    <published>2026-04-30T00:00:00+00:00</published>
  </entry>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/interactive-microwave-tla/</id>
    <title>Interactive Microwave: TLA+ in the Browser</title>
    <updated>2026-04-17T00:00:00+00:00</updated>
    <author>
      <name>Eric Spencer</name>
    </author>
    <content type="html">&lt;section id="interactive-microwave-tla-in-the-browser"&gt;

&lt;p&gt;We published a browser-native version of the
&lt;a class="reference external" href="https://github.com/EricSpencer00/interactive-microwave-tla"&gt;interactive-microwave-tla&lt;/a&gt;
simulator. It mimics a Java microwave runtime and a small TLA+ model checker
that is a literal transcription of the project’s &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;Microwave.tla&lt;/span&gt;&lt;/code&gt; spec, so
students can experiment with the state machine — including the stuttering
liveness failure from Laufer, Mertin, and Thiruvathukal,
&lt;a class="reference external" href="https://arxiv.org/abs/2407.21152"&gt;arXiv:2407.21152&lt;/a&gt; — without installing
Java, Maven, or TLC.&lt;/p&gt;
&lt;section id="the-stuttering-failure-made-visible"&gt;
&lt;h2&gt;The stuttering failure, made visible&lt;/h2&gt;
&lt;p&gt;Section IV.C of the paper introduces a liveness property and its stuttering
counterexample:&lt;/p&gt;
&lt;div class="highlight-text notranslate"&gt;&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;HeatLiveness == (radiation = ON) ~&amp;gt; (radiation = OFF)

Spec == Init /\ [][Next]_&amp;lt;&amp;lt;door, time, radiation, power&amp;gt;&amp;gt;
&lt;/pre&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Under this original &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;Spec&lt;/span&gt;&lt;/code&gt;, TLA+ permits behaviors where every step after the
microwave begins radiating is a stutter (&lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;vars'&lt;/span&gt; &lt;span class="pre"&gt;=&lt;/span&gt; &lt;span class="pre"&gt;vars&lt;/span&gt;&lt;/code&gt;). The tick never
fires, the time never decrements, and &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;radiation&lt;/span&gt; &lt;span class="pre"&gt;=&lt;/span&gt; &lt;span class="pre"&gt;ON&lt;/span&gt;&lt;/code&gt; holds forever —
violating &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;HeatLiveness&lt;/span&gt;&lt;/code&gt;. The fix (Exercise 3b) is weak fairness on &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;Tick&lt;/span&gt;&lt;/code&gt;:&lt;/p&gt;
&lt;div class="highlight-text notranslate"&gt;&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;Spec == Init /\ [][Next]_&amp;lt;&amp;lt;door, time, radiation, power&amp;gt;&amp;gt; /\ WF_vars(Tick)
&lt;/pre&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Our previous web demo did not model any of this — neither the liveness
property, nor the stuttering failure, nor the fairness condition. It merely
happened to tick unconditionally in Java and therefore looked correct without
explaining why.&lt;/p&gt;
&lt;/section&gt;
&lt;section id="what-the-new-build-does"&gt;
&lt;h2&gt;What the new build does&lt;/h2&gt;
&lt;p&gt;The browser port re-implements the spec literally in TypeScript and exposes
both the failure and its fix:&lt;/p&gt;
&lt;ul class="simple"&gt;
&lt;li&gt;&lt;p&gt;A &lt;strong&gt;Weak fairness on Tick&lt;/strong&gt; toggle in the sidebar switches the tick loop
between the paper’s original (unfair) spec and the fixed spec with
&lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;WF_vars(Tick)&lt;/span&gt;&lt;/code&gt;.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;With fairness off, the engine lets stutters run even when &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;Tick&lt;/span&gt;&lt;/code&gt; is
enabled. A runtime liveness detector surfaces the trap: after a short window
of stutter-while-radiating, the UI flags &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;HeatLiveness&lt;/span&gt;&lt;/code&gt; as violated and
shows a 🔥 overlay — a runtime witness of Figure 7.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;With fairness on, the engine forces &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;Tick&lt;/span&gt;&lt;/code&gt; to fire whenever it is
continuously enabled. &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;HeatLiveness&lt;/span&gt;&lt;/code&gt; holds and the overlay clears.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;An exhaustive safety check enumerates every state reachable from &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;Init&lt;/span&gt;&lt;/code&gt; via
the &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;Next&lt;/span&gt;&lt;/code&gt; relation and confirms zero &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;DoorSafety&lt;/span&gt;&lt;/code&gt; violations (matching
Exercise 2b).&lt;/p&gt;
&lt;/section&gt;
&lt;section id="try-it"&gt;
&lt;h2&gt;Try it&lt;/h2&gt;
&lt;div id="microwave-root" style="min-height: 560px;"&gt;&lt;/div&gt;
&lt;link rel="stylesheet" href="../../_static/demos/microwave/assets/index.css" /&gt;
&lt;script type="module" src="../../_static/demos/microwave/assets/index.js"&gt;&lt;/script&gt;&lt;p&gt;Reproduce the paper’s Figure 7 by pressing &lt;strong&gt;Power&lt;/strong&gt;, &lt;strong&gt;+3s&lt;/strong&gt;, &lt;strong&gt;Start&lt;/strong&gt;, and
then turning off &lt;strong&gt;Weak fairness on Tick&lt;/strong&gt;. The &lt;code class="docutils literal notranslate"&gt;&lt;span class="pre"&gt;Tick&lt;/span&gt;&lt;/code&gt; action stops firing,
the radiation indicator stays lit, and the liveness signal flips to
&lt;em&gt;violated&lt;/em&gt;.&lt;/p&gt;
&lt;/section&gt;
&lt;section id="source-and-details"&gt;
&lt;h2&gt;Source and details&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;a class="reference external" href="https://github.com/EricSpencer00/interactive-microwave-tla"&gt;Repository&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a class="reference external" href="https://github.com/EricSpencer00/interactive-microwave-tla/blob/main/docs/superpowers/specs/2026-04-17-web-port-design.md"&gt;Design spec&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;a class="reference external" href="https://github.com/EricSpencer00/interactive-microwave-tla/tree/main/web"&gt;Browser port (web/)&lt;/a&gt;&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;aside class="system-message"&gt;
&lt;p class="system-message-title"&gt;System Message: INFO/1 (&lt;span class="docutils literal"&gt;/home/runner/work/ai4fm/ai4fm/src/posts/interactive-microwave-tla.rst&lt;/span&gt;, line 9); &lt;em&gt;&lt;a href="#id1"&gt;backlink&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Duplicate explicit target name: “arxiv:2407.21152”.&lt;/p&gt;
&lt;/aside&gt;
&lt;p&gt;Laufer, Mertin, Thiruvathukal. &lt;em&gt;WIP: An Engaging Undergraduate Intro to
Model Checking in Software Engineering Using TLA+.&lt;/em&gt;
&lt;a class="reference external" href="https://arxiv.org/abs/2407.21152"&gt;arXiv:2407.21152&lt;/a&gt;, 2024.&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/section&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/interactive-microwave-tla/"/>
    <summary>We published a browser-native version of the
interactive-microwave-tla
simulator. It mimics a Java microwave runtime and a small TLA+ model checker
that is a literal transcription of the project’s Microwave.tla spec, so
students can experiment with the state machine — including the stuttering
liveness failure from Laufer, Mertin, and Thiruvathukal,
arXiv:2407.21152 — without installing
Java, Maven, or TLC.</summary>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="Liveness" label="Liveness"/>
    <category term="Stuttering" label="Stuttering"/>
    <category term="TLA+" label="TLA+"/>
    <published>2026-04-17T00:00:00+00:00</published>
  </entry>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/chattla-presentation-2026/</id>
    <title>Eric Spencer Presents ChatTLA+ at Loyola University Chicago</title>
    <updated>2026-04-17T00:00:00+00:00</updated>
    <author>
      <name>Arslan Bisharat</name>
    </author>
    <content type="html">&lt;section id="eric-spencer-presents-chattla-at-loyola-university-chicago"&gt;

&lt;p&gt;Eric Spencer presented &lt;strong&gt;ChatTLA+&lt;/strong&gt; at the &lt;strong&gt;Undergraduate Research and Engagement Symposium&lt;/strong&gt;
at Loyola University Chicago on April 17, 2026 - a talk exploring the use of large language
models for generating and verifying TLA+ formal specifications.&lt;/p&gt;
&lt;figure class="align-default" id="id1"&gt;
&lt;img alt="Eric Spencer presenting ChatTLA+" src="https://ericspencer.us/ai4fm/_images/eric-chattla-presentation-2026.png" style="width: 100%;" /&gt;
&lt;figcaption&gt;
&lt;p&gt;&lt;span class="caption-text"&gt;Eric Spencer presenting ChatTLA+ at the Undergraduate Research and Engagement Symposium, Loyola University Chicago, April 17, 2026.&lt;/span&gt;&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The presentation covers how LLMs can be applied to TLA+ formal specification generation
and verification, building on the group’s ongoing research into evaluating LLMs on formal
language tasks.&lt;/p&gt;
&lt;p&gt;&lt;a class="reference download internal" download="" href="../_downloads/18d709c5d4cd08bfa212db0e4661901e/ChatTLA-presentation-2026.pdf"&gt;&lt;code class="xref download docutils literal notranslate"&gt;&lt;span class="pre"&gt;View&lt;/span&gt; &lt;span class="pre"&gt;the&lt;/span&gt; &lt;span class="pre"&gt;presentation&lt;/span&gt; &lt;span class="pre"&gt;(PDF)&lt;/span&gt;&lt;/code&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a class="reference external" href="../../papers/chattla-2026/"&gt;Read the full details&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/chattla-presentation-2026/"/>
    <summary>Eric Spencer presented ChatTLA+ at the Undergraduate Research and Engagement Symposium
at Loyola University Chicago on April 17, 2026 - a talk exploring the use of large language
models for generating and verifying TLA+ formal specifications.</summary>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="LLMs" label="LLMs"/>
    <category term="Presentation" label="Presentation"/>
    <category term="TLA+" label="TLA+"/>
    <published>2026-04-17T00:00:00+00:00</published>
  </entry>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/gsirs-llm-tla-poster-2026/</id>
    <title>Brian Ortiz Presents Poster at Graduate School Interdisciplinary Research Symposium 2026</title>
    <updated>2026-04-11T00:00:00+00:00</updated>
    <author>
      <name>Arslan Bisharat</name>
    </author>
    <content type="html">&lt;section id="brian-ortiz-presents-poster-at-graduate-school-interdisciplinary-research-symposium-2026"&gt;

&lt;p&gt;Brian Ortiz presented a poster at the Graduate School Interdisciplinary Research Symposium
(GSIRS 2026) on April 11, 2026, at Loyola University Chicago.&lt;/p&gt;
&lt;p&gt;The poster, co-authored with Arslan Bisharat, Mohammed Abuhamad, Konstantin Läufer,
Eric Spencer, Khushboo Bhadauria, George K. Thiruvathukal, and TaiNing Wang, presents
the first systematic evaluation of LLM-based TLA+ specification synthesis from natural
language. The study evaluates 30 LLMs across eight families on a curated dataset of 205
TLA+ specifications, finding that LLMs achieve up to 26.6% syntactic correctness but
only 8.6% semantic correctness. Model size does not reliably predict performance on
formal languages.&lt;/p&gt;
&lt;figure class="align-default" id="id1"&gt;
&lt;img alt="Brian Ortiz at GSIRS 2026" src="https://ericspencer.us/ai4fm/_images/gsirs-2026-brian-poster.png" style="width: 100%;" /&gt;
&lt;figcaption&gt;
&lt;p&gt;&lt;span class="caption-text"&gt;Brian Ortiz standing next to his poster at the Graduate School Interdisciplinary Research Symposium 2026, Loyola University Chicago.&lt;/span&gt;&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This poster is based on the &lt;a class="reference external" href="../../papers/llm-tla-evaluation/"&gt;published ICSOFT 2026 paper&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a class="reference external" href="https://doi.org/10.6084/m9.figshare.31988706"&gt;View the poster on figshare&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a class="reference external" href="../../papers/llm-tla-evaluation/"&gt;Read the full paper details&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/gsirs-llm-tla-poster-2026/"/>
    <summary>Brian Ortiz presented a poster at the Graduate School Interdisciplinary Research Symposium
(GSIRS 2026) on April 11, 2026, at Loyola University Chicago.</summary>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="GSIRS" label="GSIRS"/>
    <category term="LLMs" label="LLMs"/>
    <category term="TLA+" label="TLA+"/>
    <published>2026-04-11T00:00:00+00:00</published>
  </entry>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/llm-tla-evaluation-2025/</id>
    <title>Can LLMs Write Correct TLA+ Specifications? Our New Evaluation Study</title>
    <updated>2026-03-29T00:00:00+00:00</updated>
    <author>
      <name>Arslan Bisharat</name>
    </author>
    <content type="html">&lt;section id="can-llms-write-correct-tla-specifications-our-new-evaluation-study"&gt;

&lt;p&gt;We have submitted a new paper evaluating whether large language models can generate
semantically correct TLA+ specifications from natural language. We evaluated 30 LLMs
across eight families on 205 TLA+ specifications. Best semantic correctness achieved
was only 8.6%, and model size did not predict quality.&lt;/p&gt;
&lt;p&gt;&lt;a class="reference external" href="../../papers/llm-tla-evaluation/"&gt;Read the full paper details&lt;/a&gt;&lt;/p&gt;
&lt;div class="admonition note"&gt;
&lt;p class="admonition-title"&gt;Note&lt;/p&gt;
&lt;p&gt;This work is now published in ICSOFT 2026. See the &lt;a class="reference external" href="../../papers/llm-tla-evaluation/"&gt;published paper&lt;/a&gt;.&lt;/p&gt;
&lt;/div&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/llm-tla-evaluation-2025/"/>
    <summary>We have submitted a new paper evaluating whether large language models can generate
semantically correct TLA+ specifications from natural language. We evaluated 30 LLMs
across eight families on 205 TLA+ specifications. Best semantic correctness achieved
was only 8.6%, and model size did not predict quality.</summary>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="LLMs" label="LLMs"/>
    <category term="TLA+" label="TLA+"/>
    <published>2026-03-29T00:00:00+00:00</published>
  </entry>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/eric-spencer-mulcahy-scholar-2025/</id>
    <title>Eric Spencer Awarded Mulcahy Scholar Stipend</title>
    <updated>2025-08-15T00:00:00+00:00</updated>
    <author>
      <name>Arslan Bisharat</name>
    </author>
    <content type="html">&lt;section id="eric-spencer-awarded-mulcahy-scholar-stipend"&gt;

&lt;p&gt;Eric Spencer, an undergraduate researcher in the AI4FM group, has been awarded a
&lt;strong&gt;$1,000 Mulcahy Scholar stipend&lt;/strong&gt; from the Loyola University Chicago College of
Arts and Sciences to support his research on LLM-based TLA+ specification synthesis.&lt;/p&gt;
&lt;p&gt;Eric is conducting this research under the supervision of Professor Konstantin Läufer.
The work contributes to the group’s ongoing evaluation of large language models for
formal verification tasks, examining whether LLMs can generate semantically correct
TLA+ specifications from natural language.&lt;/p&gt;
&lt;p&gt;The Mulcahy Scholars Program supports undergraduate research experiences of scholarly
significance across the College of Arts and Sciences at Loyola University Chicago.&lt;/p&gt;
&lt;p&gt;&lt;a class="reference external" href="https://www.luc.edu/cas/academics/undergraduateresearchopportunities/"&gt;Read more about the Mulcahy Scholars Program&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/eric-spencer-mulcahy-scholar-2025/"/>
    <summary>Eric Spencer, an undergraduate researcher in the AI4FM group, has been awarded a
$1,000 Mulcahy Scholar stipend from the Loyola University Chicago College of
Arts and Sciences to support his research on LLM-based TLA+ specification synthesis.</summary>
    <category term="Awards" label="Awards"/>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="LLMs" label="LLMs"/>
    <category term="TLA+" label="TLA+"/>
    <published>2025-08-15T00:00:00+00:00</published>
  </entry>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/gcasr-2025-poster/</id>
    <title>Brian Ortiz Presents Poster at GCASR 2025</title>
    <updated>2025-05-08T00:00:00+00:00</updated>
    <author>
      <name>Arslan Bisharat</name>
    </author>
    <content type="html">&lt;section id="brian-ortiz-presents-poster-at-gcasr-2025"&gt;

&lt;p&gt;Brian Ortiz presented a work-in-progress poster at the 12th Greater Chicago Area
Systems Research Workshop (GCASR 2025), held on May 8, 2025 at Loyola University
Chicago’s Water Tower Campus.&lt;/p&gt;
&lt;p&gt;The poster, co-authored with Mohammed Abuhamad, TaiNing Wang, George K. Thiruvathukal,
and Konstantin Läufer, presents an automated pipeline for synthesizing TLA+ specifications
from natural language using large language models, validated with the SANY parser and
TLC model checker.&lt;/p&gt;
&lt;figure class="align-default" id="id1"&gt;
&lt;img alt="Brian Ortiz and Konstantin Läufer at GCASR 2025" src="https://ericspencer.us/ai4fm/_images/gcasr-2025-brian-poster.jpeg" style="width: 100%;" /&gt;
&lt;figcaption&gt;
&lt;p&gt;&lt;span class="caption-text"&gt;Brian Ortiz standing with Konstantin Läufer next to their poster at GCASR 2025, Loyola University Chicago.&lt;/span&gt;&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;a class="reference external" href="../../papers/gcasr-2025-tla-llm/"&gt;Read the full paper details&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/gcasr-2025-poster/"/>
    <summary>Brian Ortiz presented a work-in-progress poster at the 12th Greater Chicago Area
Systems Research Workshop (GCASR 2025), held on May 8, 2025 at Loyola University
Chicago’s Water Tower Campus.</summary>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="GCASR" label="GCASR"/>
    <category term="LLMs" label="LLMs"/>
    <category term="TLA+" label="TLA+"/>
    <published>2025-05-08T00:00:00+00:00</published>
  </entry>
  <entry>
    <id>https://ericspencer.us/ai4fm/posts/tla-for-all-running-model-checking-in-a-python-notebook/</id>
    <title>TLA+ for All: Running Model Checking in a Python Notebook</title>
    <updated>2025-02-05T00:00:00+00:00</updated>
    <author>
      <name>Arslan Bisharat</name>
    </author>
    <content type="html">&lt;section id="tla-for-all-running-model-checking-in-a-python-notebook"&gt;

&lt;p&gt;We published a project embedding TLA+ model checking directly into a Python notebook
environment, making formal verification accessible without any installation. Users can
write specifications, run the TLC model checker, and visualize behaviors in-browser.&lt;/p&gt;
&lt;p&gt;&lt;a class="reference external" href="../../papers/tla-for-all/"&gt;Read the full paper details&lt;/a&gt;&lt;/p&gt;
&lt;/section&gt;
</content>
    <link href="https://ericspencer.us/ai4fm/posts/tla-for-all-running-model-checking-in-a-python-notebook/"/>
    <summary>We published a project embedding TLA+ model checking directly into a Python notebook
environment, making formal verification accessible without any installation. Users can
write specifications, run the TLC model checker, and visualize behaviors in-browser.</summary>
    <category term="FormalMethods" label="Formal Methods"/>
    <category term="TLA+" label="TLA+"/>
    <published>2025-02-05T00:00:00+00:00</published>
  </entry>
</feed>
