Projects
Configurable BCH Decoder for High-Speed Optical Interconnects
- Architected a multi-standard BCH decoder in Verilog supporting (63, 51)/(255, 239)/(1023, 983) codes in hard/soft-decision modes, optimized via a unified GF multiplier with XOR-tree reduction, 8-lane shared ROM, and correlation metric.
- Implemented the full RTL-to-GDSII flow using VCS and Innovus, verifying functional correctness and timing closure via simulation and post-layout STA, delivering 1.58 mm² core area, 14.7 mW power, and zero DRC/LVS violations.
AT²-Optimized 32-bit Pipelined RISC-V Processor (RV32IC)
- Designed a 32-bit, 5-stage pipelined RISC-V processor (RV32I + C extension) in Verilog with 2-way set-associative I/D caches, a hazard unit for data/control hazards via forwarding/stalls, and early branch resolution halving misprediction penalty.
- Compared branch prediction and multiplier architectures via AT² analysis; synthesized in TSMC 0.13µm achieving 3.0 ns cycle time and 3.17×10⁵ µm² area, with C-extension compression cutting execution time up to 37%.
Incremental SAT-PBO ATPG Framework on Berkeley ABC
- Developed an incremental SAT+PBO-based ATPG framework in Berkeley ABC, adding custom commands for fault modeling and equivalence fault collapsing, with a PBO objective maximizing faults detected per pattern to minimize test set size.
- Validated the framework on the c17 ISCAS85 benchmark, achieving fewer test patterns than PODEM with 100% fault coverage and diagnosis accuracy via the Kissat SAT solver, while identifying clause-explosion scalability limits on larger designs.
A2C-Based Logic Optimization under Black-Box Cost Functions
- Built an RL-based logic synthesis optimizer using the A2C algorithm on Yosys-ABC, formulating command-sequence selection as a POMDP, with library preprocessing selecting lowest-cost gate variants for technology mapping.
- Evaluated across 48 test cases, benchmarking against Greedy, Simulated Annealing, and Fast Simulated Annealing baselines, with RL achieving up to 55.3% lower cost than Greedy and remaining the most consistent performer overall.
Image reference: ICCAD Contest 2024
T-Count Optimization Framework for Clifford+T Quantum Circuits
- Developed a unified T-count optimization framework that integrates multiple reduction techniques including TMerge, Internal-H-OPT and advanced phase polynomial methods such as TODD for efficient circuit synthesis.
- Enhanced overall circuit efficiency by combining Gray synthesis (GraySyn) with T-parallelism (T-Par) strategies to exploit structural regularities and gate-level concurrency.
Quantum Circuit Enumeration with Unitary Matrix
- Read in a valid unitary matrix, converting the tensor into several 2-level matrices.
- Use gray-code synthesis to map the matrices into quantum gates, decomposing and optimizing to get the final quantum circuit with the given basic gate sets.
NTUEE LightDance
- Established a C/C++ library for the communication between ATTiny85 board and RPI.
- Constructed the full data path from the latch to the microcontroller and then to the Raspberry Pi for WS2812 LED strips, and from the PCA9955 IC to the Raspberry Pi for optical fiber.
Multimodal Perception of Corner Cases in Autonomous Driving
- Developed a system for multimodal perception and comprehension in autonomous driving, focusing on global scene understanding, local area reasoning, and actionable navigation using the CODA-LM dataset.
- Enhanced the perception capabilities of LLaVA 1.5 7b by fine-tuning LoRA and incorporating additional modules to handle diverse scenes, small objects, and complex driving scenarios effectively.
Image reference: CODA 2024 Workshop
Smart Meeting Cube: Embedded System for Interactive Meeting Management
- Designed a Smart Meeting Cube that streamlines meeting management tasks such as attendance, voting, and speaking requests through interactive cube rotations.
- Implemented face recognition and real-time communication between STM32 and Raspberry Pi using WebSocket.