ucsc-ospo
diff --git a/‎content/authors/YimingCheng/_index.md
Lines changed: 53 additions & 0 deletions b/‎content/authors/YimingCheng/_index.md
Lines changed: 53 additions & 0 deletions
diff --git a/‎content/authors/YimingCheng/avatar.jpg
200 KB b/‎content/authors/YimingCheng/avatar.jpg
200 KB
diff --git a/‎content/project/osre25/uchicago/EnvGym/featured.png
1.73 MB b/‎content/project/osre25/uchicago/EnvGym/featured.png
1.73 MB
diff --git a/‎content/project/osre25/uchicago/EnvGym/index.md
Lines changed: 54 additions & 0 deletions b/‎content/project/osre25/uchicago/EnvGym/index.md
Lines changed: 54 additions & 0 deletions
diff --git a/‎content/report/osre25/uchicago/EnvGym/featured.png
1.68 MB b/‎content/report/osre25/uchicago/EnvGym/featured.png
1.68 MB
diff --git a/‎content/report/osre25/uchicago/EnvGym/index.md
Lines changed: 56 additions & 0 deletions b/‎content/report/osre25/uchicago/EnvGym/index.md
Lines changed: 56 additions & 0 deletions
@@ -0,0 +1,53 @@
+---
+# Display name
+title: Yiming Cheng
+
+# Username (this should match the folder name)
+authors:
+  - YimingCheng
+
+# Is this the primary user of the site?
+superuser: false
+
+# Role/position
+role: "Predoc Researcher, Department of Computer Science, University of Chicago"
+
+# Organizations/Affiliations
+organizations:
+  - name: University of Chicago
+    url: "https://computerscience.uchicago.edu/"
+
+# Short bio (displayed in user profile at end of posts)
+bio: Yiming Cheng is a Pre-doc researcher in the Department of Computer Science at the University of Chicago, pursuing an M.S. in Computer Science with a focus on Machine Learning Systems (MLSys). He graduated from Tsinghua University with a B.E. in Electronic Engineering. His research focuses on distributed LLM deployment, distributed KV cache, and efficient machine learning systems. He is currently working on the LMCache project and contributing to open-source initiatives in the machine learning systems space.
+
+# Social/Academic Networking
+# For available icons, see: https://sourcethemes.com/academic/docs/widgets/#icons
+#   For an email link, use "fas" icon pack, "envelope" icon, and a link in the
+#   form "mailto:your-email@example.com" or "#contact" for contact widget.
+social:
+  - icon: envelope
+    icon_pack: fas
+    link: mailto:eaminc0328@gmail.com
+  - icon: globe
+    icon_pack: fas
+    link: https://eaminc.github.io/
+
+# Enter email to display Gravatar (if Gravatar enabled in Config)
+email: ""
+
+# Organizational groups that you belong to (for People widget)
+#   Set this to `[]` or comment out if you are not using People widget.
+user_groups:
+  - Summer of Reproducibility Mentors
+  - 2025 Contributors
+---
+
+Yiming Cheng is a Pre-doc researcher in the Department of Computer Science at the University of Chicago, where he is pursuing an M.S. in Computer Science with a specialization in Machine Learning Systems (MLSys track). He graduated from Tsinghua University in 2024 with a Bachelor of Engineering in Electronic Engineering, along with minors in Statistics and Law.
+
+Currently, Yiming is working as an Open Source Contributor and Research Assistant with the LMCache team under the supervision of Prof. Junchen Jiang. His work focuses on LMCache, the first open-source Knowledge Delivery Network (KDN) that accelerates LLM applications up to 8x faster at 8x lower cost. He also contributes to vLLM/production-stack, helping scale from single vLLM instances to distributed vLLM deployments. He has contributed over 1,262 lines of code to these open-source projects.
+
+His research interests span both systems for machine learning (distributed LLM deployment, distributed KV cache, efficient ML) and machine learning for systems (ML for code generation and Operating Systems). During his undergraduate studies, he worked extensively on data mining projects including recommendation systems, emotion awareness, and embodied city simulations.
+
+Yiming has been recognized with several prestigious awards, including the Merit-based Predoc Scholarship of $40,000 from the University of Chicago and funding from the United States National Science Foundation for his Summer of Reproducibility (SoR) project. He has authored multiple publications in venues such as MDPI Sensors and has patents in semantic encoding and decoding frameworks.
+
+Through his research at institutions including Argonne National Laboratory, Tsinghua University's Future Intelligent Lab, and the University of Houston, Yiming continues to contribute to advancing distributed computing technologies and machine learning systems. His expertise in programming languages including Python (PyTorch, CuPy), Go (Docker, K8s), and various other technologies makes him a valuable contributor to the open-source machine learning community.
@@ -0,0 +1,54 @@
+---
+title: "Smart Environments – An AI System for Reproducible Custom Computing Environments"
+authors: [marshalp, YimingCheng]
+author_notes: ["University of Chicago", "University of Chicago"]
+tags: ["osre25", "reproducibility", "machine learning", "OS"]
+date: 2025-02-18
+lastmod: 2025-02-18
+---
+
+## Overview
+
+The complexity of environment setup and the expertise required to configure specialized software stacks can often hinder efforts to reproduce important scientific achievements in HPC and systems studies. Researchers often struggle with incomplete or ambiguous artifact descriptions that make assumptions about "common knowledge" that is actually specific domain expertise. When trying to reproduce experiments, reviewers may spend excessive time debugging environment inconsistencies rather than evaluating the actual research. These challenges are compounded when experiments need to run on different hardware configurations.
+
+This project seeks to address these fundamental reproducibility barriers by using AI to translate natural language environment requirements often used in papers or artifact descriptions into actionable, reproducible configurations—bridging the knowledge gap between experiment authors and reviewers while standardizing environment creation across different hardware platforms. We will develop an AI-driven system that automatically generates and configures reproducible computing environments based on artifact descriptions from conferences, Trovi artifacts on the [Chameleon](chameleoncloud.org) testbed, and other reliable sources for scientific experiment code and associated documentation. Leveraging Natural Language Processing (NLP), the system will allow researchers to describe desired environments in plain English, then map those descriptions onto predefined configuration templates. By simplifying environment creation and ensuring reproducibility, the system promises to eliminate duplicate setup efforts, accelerate research workflows, and promote consistent experimentation practices across diverse hardware.
+
+## Key Outcomes
+
+- Working Prototype: A system that automatically generates machine images deployable on bare metal and VM instances, based on user-provided requirements.
+- Comprehensive Documentation: Detailed user manuals, guides, and best practices tailored to researchers, ensuring a smooth adoption process.
+- Live Demo: A demonstration environment (e.g., a web app or Jupyter notebook) that shows how to request, configure, and launch reproducible cloud environments on both hardware profiles.
+- Long-Term Impact: Building blocks for future AI-driven automation of cloud infrastructure, reducing human error and enabling fast, repeatable research pipelines.
+
+**Topics**: Reproducibility, AI & NLP, Cloud Computing, DevOps and Automation
+
+**Skills**:
+
+- Machine Learning / AI: Familiarity with NLP methods to interpret user requirements.
+- Python: Primary language for backend services and cloud interactions.
+- Cloud API Integration: Experience with OpenStack or similar APIs to provision and configure images on both bare metal and virtual machines.
+- DevOps: Automated environment configuration, CI/CD workflows, and containerization.
+
+**Difficulty**: Hard
+
+**Size**: Large
+
+**Mentors**: {{% mention marshalp %}}
+
+**Tasks**:
+
+- Requirement Gathering & NLP Design
+  - Research the specific needs of researchers building experimental setups.
+  - Design an NLP pipeline to parse plain-English descriptions (e.g., “I need Python 3.9, CUDA 11, and scikit-learn”) into environment “recipes.”
+- Backend Environment Builder
+  - Implement logic that converts parsed user requirements into machine-image definitions for bare metal and VM instances.
+  - Integrate with Chameleon’s APIs to provision servers, install software, and run configuration validation automatically.
+- Front-End & User Experience
+  - Develop an intuitive web or CLI interface that researchers can use to capture experiment environment requirements.
+  - Provide real-time status updates during environment setup, along with meaningful error messages and quick-start templates.
+- Testing & Validation
+  - Conduct end-to-end tests using diverse software stacks (e.g., HPC libraries, machine learning frameworks) on bare metal and VM instances.
+  - Ensure reproducibility by re-creating the same environment multiple times and comparing configurations.
+- Documentation & Demonstration
+  - Produce user-facing documentation, including tutorials and best practices for researchers who frequently run experiments on Chameleon Cloud.
+  - Create a short live demo or screencast showcasing how to configure an environment for a specific research workflow.
@@ -0,0 +1,56 @@
+---
+title: "EnvGym – An AI System for Reproducible Custom Computing Environments"
+subtitle: ""
+summary:
+authors:
+  - YimingCheng
+tags: ["osre25", "reproducibility", "machine learning", "OS"]
+categories: ["osre25", "reproducibility", "EnvGym"]
+date: 2025-06-16
+lastmod: 2025-06-16
+featured: false
+draft: false
+
+# Featured image
+# To use, add an image named `featured.jpg/png` to your page's folder.
+# Focal points: Smart, Center, TopLeft, Top, TopRight, Left, Right, BottomLeft, Bottom, BottomRight.
+image:
+  caption: "EnvGym Project"
+  focal_point: Top
+  preview_only: false
+---
+
+Hello, My name is Yiming Cheng. I am a Pre-doc researcher in Computer Science at University of Chicago. I'm excited to be working with the Summer of Reproducibility and the Chameleon Cloud community as a project leader. My project is [EnvGym](https://github.com/eaminc/envgym) that focuses on developing an AI-driven system to automatically generate and configure reproducible computing environments based on natural language descriptions from artifact descriptions, Trovi artifacts, and research papers.
+
+The complexity of environment setup often hinders reproducibility in scientific computing. My project aims to bridge the knowledge gap between experiment authors and reviewers by translating natural language requirements into actionable, reproducible configurations using AI and NLP techniques.
+
+### Project Overview
+
+EnvGym addresses fundamental reproducibility barriers by:
+
+- Using AI to translate natural language environment requirements into actionable configurations
+- Automatically generating machine images deployable on bare metal and VM instances
+- Bridging the knowledge gap between experiment authors and reviewers
+- Standardizing environment creation across different hardware platforms
+
+### June 10 – June 16, 2025
+
+Getting started with the project setup and initial development:
+
+- I began designing the NLP pipeline architecture to parse plain-English descriptions (e.g., "I need Python 3.9, CUDA 11, and scikit-learn") into structured environment "recipes"
+- I set up the initial project repository and development environment
+- I met with my mentor Prof. Kexin Pei to discuss the project roadmap and technical approach
+- I started researching existing artifact descriptions from conferences and Trovi to understand common patterns in environment requirements
+- I began prototyping the backend environment builder logic that will convert parsed requirements into machine-image definitions
+- I explored Chameleon's APIs for provisioning servers and automated configuration
+
+### Next Steps
+
+- Continue developing the NLP component for requirement parsing
+- Implement the core backend logic for environment generation
+- Begin integration with Chameleon Cloud APIs
+- Start building the user interface for environment specification
+
+This is an exciting and challenging project that combines my interests in AI systems and reproducible research. I'm looking forward to building a system that will help researchers focus on their science rather than struggling with environment setup issues.
+
+Thanks for reading, I will keep you updated as I make progress on EnvGym!