Evaluating Large Language Models Trained on Code
作者:Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harrison Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Łukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth A. Barnes, Ariel Herbert-Voss, William H. Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, I. Babuschkin, Suchir Balaji, Shantanu Jain, William S. Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew M. Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, Wojciech Zaremba · 发表于:arXiv (Cornell University) · 年份:2021 · DOI:10.48550/arxiv.2107.03374 · 被引用次数:1458 · 研究领域:Software Engineering Research、Parallel Computing and Optimization Techniques、Topic Modeling
I created a basic proof of concept of LMVM (Language Model Virtual Machine), a command line toolchain that English instructions into native binaries without intermediary high-level language compilation. The command line uses a large language model (LLM) as IR generation backend, producing LLVM Intermediate Representation (IR) directly from user instructions coding. The IR is compiled and linked to machine code (-o files) via `clang` (with an `llc`/`lld` fallback on POSIX hosts) and executed on bare metal, within Docker sandboxes, on AWS Lambda, or as standalone HTTP servers. How It Works (5 Steps) 1. You type an instruction - e.g. "Create a rest api running on localhost port 3000 and return hello world" 2. LMVM checks your computer - It figures out if you're on Windows/Mac/Linux, what CPU you have (Intel vs ARM), and makes sure clang (the compiler) is installed. If not, it offers to install it for you (with your permission). 3. An AI writes LLVM IR - Instead of writing C or Python, the LLM writes LLVM IR -- a low-level assembly-like language that sits between source code and machine code. The prompt includes your OS details so the AI generates code that works on your platform (e.g. Winsock on Windows vs POSIX sockets on Mac/Linux). 4. Clang turns IR into a native binary - clang compiles the IR and links it with system libraries (PostgreSQL, OpenSSL, SQLite, etc.) into a real executable. One step, no extra tools needed. 5. It runs - The binary executes directly on your CPU. If...