Day 1 · R2 Platform Services

Compute-platform and hardware leaders on AI compute, chip strategy, and enterprise adoption playbooks.

Day 1 · Fri 27 Sep 2024Room 2 — Platform Services13:00–17:00Kao Hung-yu (高宏宇), Professor, College of Informatics, National Tsing Hua University

At a glance

SpeakerTalkIn one sentence
Tsai Chi-yen (蔡祈岩), Deputy General Manager and Chief Information Officer, Taiwan MobileAI, the Age of ItUsing an 'Age of It' view of technology history, he proposes a Five Dimensions framework for enterprise AI adoption plus a Taiwan Mobile customer-service case.
Hsu Ta-yung (徐達勇), Director of Applications Engineering, ArmBuilding the AI Compute Platform of the FutureCentered on 'Change is Relentless' and 'Multi-Die Everywhere', he cites AI-potential data across healthcare, automotive, education, and climate change.
Chang Ou-yu-hao (張歐佑豪), Senior Technical Consultant, AMD Taiwan Commercial Business DivisionFrom Silicon Design to Software Stack: AMD AI StrategyHe argues for 'designing different architectures for different AI needs', using the MI300X's 192GB HBM3 versus the H100 to show fewer GPUs are needed to run a model.
Yang Tzu-hsiang (楊子翔), Senior Director, Quanta ComputerFrom no-code, low-code to full-code AI model DevOps: QOCA aimQuanta's no-code medical AI platform QOCA aim completes modeling within 2 hours and has already partnered with multiple hospitals and international workshops.
Lin Wei (林緯), Chief Technology Officer, Phison ElectronicsBreaking Barriers in Generative AI Deployment: Phison aiDAPTIVPhison uses NAND Flash to expand GPU 'memory' rather than compute, cutting the cost of training a 90B model from NT$30 million to under NT$1 million.

Talks

5 of 5 talks

1AI, the Age of It: A Great Opportunity for Taiwanese BusinessesTsai Chi-yen (蔡祈岩), Deputy General Manager and Chief Information Officer, Taiwan Mobile

Using an 80-year 'one mega wave every 20 years' view of IT history, he proposes a Five Dimensions framework for enterprise AI adoption plus a Taiwan Mobile customer-service case.

Key points

  • In his opening self-introduction, he mentioned that as a freshman in Computer Science at National Chiao Tung University he made a game called 'Decisive Battle Russia' (決戰俄羅斯), which sold about 100,000 copies in Taiwan; the royalties paid for his university and graduate-school tuition — this fully matches public reporting (see verification notes).
  • He proposes a view of IT history with 'one Mega Wave every 20 years': 1946 — governments use mainframes (control of the postwar world order) → 1959 — the IBM 7090 lets enterprises own computers (MRP/ERP/global logistics) → 1977–81 — the Apple/IBM PC brings computers to the desktop, giving rise to the knowledge economy and the Internet (Taiwan used this to enter IBM's open-architecture supply chain, the starting point of today's prosperity of TSMC and the semiconductor industry) → 1999 — PALM kicks off mobile devices (mainland China was the biggest beneficiary; Taiwan only grew by inertia) → 2020–2040 — AI (which he names the 'Age of It', where 'it' — as opposed to 'he/she' — represents a being capable of thought).
  • He argues two signs show the 'Age of It' has already begun: (1) the Stanford 2024 AI Index shows that after 2017/18, industry's number of large-model investments crossed over academia's in a 'golden cross'; (2) AI performance benchmark curves keep approaching and even surpassing the human baseline.
  • He proposes a 'Five Dimensions' framework for enterprise AI adoption (any two dimensions can intersect to form a strategic plane, and any three can form a strategic space): ① raising employee productivity ② connecting digital platforms (internal ERP/HR/finance systems as well as external websites/apps can all have an 'AI brain installed') ③ linking into the AI ecosystem (open-source communities, Hugging Face — he also argues this improves talent retention) ④ training the enterprise's own models ⑤ AI ethics (information security, copyright, legal compliance — which he considers, at this stage, a technically controllable problem).
  • Two internal mechanisms drive AI adoption: Inside-out — an AI Hackathon is held on the last Friday of every month, where IT engineers may only pitch ideas for systems they are currently developing or maintaining themselves (they may not propose ideas for someone else's system); each pitch runs 3–5 minutes to 5 IT division heads and needs 3 approval votes to pass. In the first year, roughly 30-plus systems passed, but only about 5 were actually adopted by business units (BUs). Outside-in — a two-day, one-night hackathon where BUs lead teams to pitch proposals; about 40 teams reach the finals, and the CEO/CXOs judge and decide resource allocation. The core principle is 'let whoever hears the gunfire give the order' — tactical decisions are made at the front line, while the CEO only decides whether to commit resources at the strategic level.
  • He emphasizes the necessity of local models: ChatGPT cannot understand the way Taiwanese people mix Mandarin, Taiwanese (Hokkien), and Hakka when speaking; in Taiwan Mobile's self-built customer-service system, the Whisper model kept mishearing '摸乳 (MO, mobile number)' as other words — a concrete case for why a localized speech model is indispensable.
  • Taiwan Mobile ultimately chose the customer-service domain to train its own model. Features already live include: the AI directly handling customer interactions, identity voice verification, voice-quality inspection of customer-service staff, call transcription, and post-call summaries. Also planned are an AI Agent (like 'a ChatGPT that understands the company's processes', able to connect to internal systems to handle leave requests and schedule meetings) and an AI Mentor (a coding mentor to train junior engineers, and a customer-service practice partner that simulates customer calls and scores agents in real time).
  • He notes that after IT engineers adopted AI, internal surveys found coding productivity rose 'at least 3x'; AI code review and translating COBOL/stored procedures into Java reached about 80–90% completion.

Tech, products & figures

Stanford 2024 AI Index
Data source for the academia/industry golden cross in large-model investment, and for AI performance approaching the human baseline.
ChatGPT
Hugging Face
Whisper
Speech-recognition model; case of misrecognition from Taiwanese/Mandarin code-mixed speech.
momo (Taiwan Mobile subsidiary)
Bet on the app early rather than the desktop browser — used to illustrate 'capturing the mobile-device dividend early'.

Notable quotes

The AI you're looking at in 2024, and you're still laughing at its hallucinations... just think back to 1984, and the toy you saw then.
Money is the most rational investor... you don't release the hawk until you see the rabbit. When you notice all the hawks have flown out, you know the old hunter must have seen a lot of rabbits.

Q&A

  • No Q&A record.

Fact-check notes

  • Verified via WebSearch: Tsai Chi-yen currently serves as Deputy General Manager and CIO of Taiwan Mobile, and previously served as CTO of DBS Bank (Taiwan), CIO of the Sinyi Realty Group, and Senior Vice President of HSBC (China). He is indeed the same person who, as a freshman, developed the game 'Decisive Battle Russia' and won the National Game Design Golden Disk Award — fully consistent with the self-introduction in his talk.Sources:經理人數位時代工商時報
2Building the AI Compute Platform of the FutureHsu Ta-yung (徐達勇), Director of Applications Engineering, Arm

Centered on 'Change is Relentless' and 'Multi-Die Everywhere', he uses data from four industries to illustrate AI's potential.

Key points

  • The core theses are 'Change is Relentless' (change never stops) and 'Multi-Die Everywhere' (multi-die design is everywhere).
  • He observes that 'design and manufacturing cycle times keep getting longer', arguing that longer cycles push up costs, which in turn drives the need to bring AI into chip design and manufacturing processes.
  • Under the heading 'AI – The Potential', he cites AI-potential data across four industries:
    • Healthcare: AI can shorten R&D cycles by 50% (original: '50% Shortened R&D Cycle'), corresponding to shorter drug-development time.
    • Automotive: AI can save over 1,000,000 additional lives each year (original: '1,000,000+ Lives Saved Each Year'), corresponding to reduced traffic accidents and fatality rates.
    • Education: AI can increase access to one-on-one tutoring by 50x (original: '50x Increase in Tutoring Access'), corresponding to personalized curriculum design.
    • Climate Change: AI can analyze/compute 10,000x faster than a human (original: '10,000x Faster than a Human'), corresponding to time saved on human analysis and computation.

Tech, products & figures

"AI – The Potential" four industry data points
Healthcare (R&D cycle −50%), Automotive (1,000,000+ additional lives saved per year), Education (tutoring access ×50), Climate Change (analysis speed ×10,000).

Notable quotes

"Change is Relentless"
"Multi-Die Everywhere"

Q&A

  • No Q&A record.

Fact-check notes

  • This section is compiled from attendees' collaborative community notes.
  • No independent corroborating source — official Arm materials or third-party — was found for these exact four 'AI potential' figures; flagged as based only on the crowd-sourced notes with no additional public source verified. Please do not extrapolate further from these figures.
3From Silicon Design to Software Stack: AMD AI StrategyChang Ou-yu-hao (張歐佑豪), Senior Technical Consultant, AMD Taiwan Commercial Business Division

He argues for 'designing different architectures for different AI needs', using the MI300X's 192GB HBM3 versus the H100 to show fewer GPUs are needed to run a model.

Key points

  • Core thesis: 'design different architectures for different AI needs' — the need types listed are general process, inference, training, (one term was hard to make out and is provisionally recorded as 'graining', possibly a slip or abbreviation — the correct term could not be confirmed), and low power.
  • Another core thesis: 'packaging technology is ahead of the market.'
  • Concrete spec comparison: MI300X vs. H100 — the MI300X has 192GB of HBM3, letting you 'run a model on fewer GPUs, without having to buy so many cards' — i.e., a larger per-card memory capacity reduces the need for multi-card parallelism.

Tech, products & figures

MI300X
AMD data-center GPU equipped with 192GB HBM3.
H100
NVIDIA data-center GPU, used as the comparison baseline.

Q&A

  • No Q&A record.

Fact-check notes

  • This section is compiled from attendees' collaborative community notes.
  • Cross-verified via WebSearch: the MI300X's official specs are indeed 192GB HBM3 with 5.3 TB/s memory bandwidth; the NVIDIA H100 is 80GB HBM3 with 3.35 TB/s bandwidth. The MI300X's memory capacity is about 2.4–2.7x that of the H100, and its bandwidth about 1.6x — matching the direction and scale of the comparison in the crowd-sourced notes.Sources:AMD 官網產品頁Lenovo Press 產品指南TRG DatacentersGPUPerHour
  • Found a concrete case corroborating 'running a model on fewer GPUs': a 70B-parameter model at FP16 precision (requiring about 140GB) fits entirely on a single MI300X (192GB), whereas on the H100 (80GB) it requires at least two cards together with tensor parallelism to run.Sources:TRG Datacenters
  • 'Packaging technology is ahead of the market': no specific public figures from AMD were found to substantiate this; flagged as based only on the crowd-sourced notes, with no additional public spec figures verified.
4From no-code, low-code to full-code AI model DevOps: QOCA aimYang Tzu-hsiang (楊子翔), Senior Director, Quanta Computer

Quanta uses its no-code platform QOCA aim to compress hospital-side AI modeling time from two weeks to two hours, while gradually adding LLM/RAG capabilities.

Key points

  • He opens with four pop-culture examples (ChatGPT, the controversy over an AI-generated painting winning an award, the film 'A.I. Artificial Intelligence', and '2001: A Space Odyssey') to make the point that AI actually has a long history and didn't just appear in 2022, paralleling the history of information technology (abacus → mainframe → PC → WWW → 2022 LLMs) with the history of medical informatics (in 1971 a hospital terminal mainframe had only 1MB of memory; Johns Hopkins' 1980 in-hospital network architecture; a 2004 paper predicting the curve of the shift from paper-based to computer-based records).
  • Industry observation: the accuracy of medical AI image recognition began rising rapidly around 2018, driving a clear increase from 2018 onward in the number of FDA/TFDA and 510(k)-cleared software, concentrated in imaging applications such as radiology (chest X-ray, CT/MRI), cardiovascular, and hematology.
  • Quanta's internal survey cites pain points such as 'the biggest barrier to AI adoption is a lack of talent (63%)' and 'data security is the top priority, with a preference for private cloud', which form the rationale for launching the no-code platform.
  • QOCA aim (Quanta Open Care AI, named after the Australian quokka, billed as 'the happiest animal in the world'):
    • 1.0: no-code modeling for structured data/medical imaging (classification, including explainability heatmaps, e.g. visualizing the hotspots behind a cardiac-image reading)
    • 2.0: expanded toward LLM/RAG/fine-tuning, also following a no-code approach, including an LLM playground for inference
    • The interface packs 20-plus function buttons on a single page, with 150+ sub-functions once fully expanded; it supports step-by-step forward/back and an overall reset for every step, and new data can be replayed in full
  • Real-world impact: after hospital-side data scientists and clinicians discuss the data, the no-code platform can produce model results (including accuracy and data-imbalance warnings) within two hours, versus the two weeks of back-and-forth discussion typical of outsourced modeling in the past. The platform has been used to teach at international workshops for two consecutive years (split into morning and evening sessions due to time-zone differences), with over 80 successful training runs and a success rate above 90%; demo data has covered cervical-cancer classification, breast imaging, and more.
  • Collaboration with MIT: an extension of MIT's Segment Anything model is borrowed for annotation assistance; a smart chest X-ray module; a radiotherapy module; the 'V-Line' mammography risk-assessment module (predicting risk over the next 5 years); and the 'SEEBOR' lung-cancer risk model (targeting lung-cancer causes among Taiwan's non-smoking population, such as air pollution and high-heat wok cooking fumes, published in the journal JCO).
  • He presents a data-lifecycle perspective: the traditional modeling workflow (EDA, looking at distribution plots, dropping missing values, setting features/target, running random forest/XGBoost) is actually just a small piece of the whole data-collection–training–serving loop; QOCA aim additionally wraps in the 'serving' piece, so a trained model can directly generate a RESTful endpoint or a web interface for input and prediction.
  • Quanta's hardware business rests on three pillars: Mobile Computing, Cloud Computing, and AI Computing; he also mentioned that Quanta has manufactured the first-generation Apple Watch, AI servers (hinting at cooperation with Nvidia/server partners), and self-driving-car components.
  • Strategic direction: the platform is currently used only in the healthcare industry, with the goal of gradually expanding to other industries (echoing the conference theme 'AI for Every Industry').

Tech, products & figures

QOCA aim (Quanta Open Care AI)
Quanta's medical AI no-code/low-code cloud platform; 1.0 focuses on imaging/structured-data modeling, 2.0 expands into LLM/RAG.
MIT Segment Anything
Extended and applied for annotation assistance.
V-Line
Mammography 5-year risk-assessment module.
SEEBOR
Lung-cancer risk model (published in JCO).

Notable quotes

"Google told us that what we just did is only that little black box in the middle of the (data lifecycle) diagram." (illustrating that modeling is just a small segment of the whole data-engineering pipeline)

Q&A

  • No Q&A record.

Fact-check notes

  • Verified via WebSearch: QOCA aim is indeed Quanta's medical AI cloud platform, offering no-code/low-code tools to help hospitals analyze medical imaging and structured data; it has already partnered with hospitals such as Taipei Medical University, Taipei Veterans General Hospital, Cathay General Hospital, and Fu Jen Catholic University Hospital (e.g. Cathay General Hospital's cardiovascular-center smart-ward system), consistent with the talk's content.Sources:數位時代成大醫院 QOCA AIM 簡介QOCA 官網
5Breaking Barriers in Generative AI Deployment: Phison aiDAPTIVLin Wei (林緯), Chief Technology Officer, Phison Electronics

Phison uses SSD/NAND Flash to expand GPU 'memory' rather than compute, cutting the on-premises training cost of a 90B model to under a tenth.

Key points

  • Company background: Phison is a NAND Flash solutions company with roughly 20% global market share (of about 3.5 billion NAND-Flash-containing electronic components shipped each year, Phison supplies about 700 million); it took part in NASA's 2021 Mars 'Perseverance' mission (the largest IC on the rover's motherboard was a Phison SSD), has recently sent SSDs to the space station and retrieved them for testing, and mentioned a concept for a 'lunar data center' in cooperation with NASA.
  • Core problem diagnosis: under the traditional AI server architecture, compute (the GPU IC) and memory (HBM) are bundled together, with NVIDIA dominating the definition of the entire software/hardware stack (from the driver down to the hardware layer), leaving the CPU, system DRAM, and traditional SSDs almost idle during training/inference (the speaker jokingly calls them 'paid slackers'). Because today's deep-learning models are 100–1,000x larger than past CV/speech models, the training bottleneck has shifted from 'not enough compute' to 'not enough VRAM/HBM' — yet NVIDIA sells compute and memory bundled together, which raises the barrier for enterprises building their own AI infrastructure.
  • Concrete cost figures: an H100 card costs nearly NT$1 million in Taiwan (rising to about NT$2 million if carried out of the country in person); HBM costs about US$200 per GB versus about US$0.1 per GB for NAND Flash (nearly a 2,000x price gap); training the LLaMA 3.2 90B multimodal model with a traditional architecture requires a build cost of at least NT$30 million; a high-speed network switch (NVLink, 400G) costs about NT$10 million per unit.
  • The aiDAPTIV solution: through in-house-developed memory-management software, the GPU can use system DRAM and NAND Flash SSDs — in addition to HBM — as a usable 'AI memory expansion layer', trading some speed for affordable capacity. They claim this approach can cut the cost of training a 90B model to under a tenth (under NT$1 million), with some customers even completing full-parameter training of a 90B model on a roughly NT$100,000 desktop PC, taking about 4 hours per 10M tokens.
  • Cluster network architecture innovation: because every node holds the complete data, there is no need to rely on high-speed 400G interconnects like NVLink as traditional multi-card training does (switches costing about NT$10 million each); ordinary 10G networking suffices instead (switches costing about NT$1,000 each). In measured tests, four machines linked over a low-speed network achieved about 4x the throughput of a single machine, with a scaling factor of 0.9 or higher.
  • Product tiers: a single-card machine (aimed at school education; a single card alone can do inference and full-parameter training for models like 8B/13B, and four linked together can train 70B/90B models with full parameters); a workstation (for small businesses); and an 8-card machine (for medium and large enterprises — Phison itself uses two 8-card machines to serve an AI-assistant system for its entire 4,000-employee workforce).
  • Three reasons for on-premises deployment: ① security — the enterprise's confidential emails/documents never leave for the cloud ② autonomy and controllability — cloud-based large models will refuse to generate certain content due to policy/sensitive-word filters (a live example: when cooperating with police on an AI transcript system, sensitive terms in child sexual exploitation cases get directly deleted and refused by ChatGPT, so police currently have to manually substitute sensitive words — e.g. calling it '外送員 (delivery rider)' instead of '車手 (money mule)' — to get around the filter) ③ cost grows exponentially with model scale, making it hard to keep expanding the budget indefinitely.
  • Internal AI assistant 'Javis': trained on the company's internal email corpus, with features including email summarization, translation, auto-generated reply drafts, attachment analysis, and turning meeting recordings into minutes; it can also auto-generate training sets (feed in documents and it automatically produces the data needed for RAG/fine-tuning), and they claim 'on-premises model + RAG + fine-tuning' outperforms 'cloud ChatGPT + RAG'.
  • Education outreach: donating adaptive-drive hardware and complete course materials/lab designs to universities (such as National Yang Ming Chiao Tung University) to help set up AI computer classrooms; also donating to police departments, hospitals, and government agencies for total solutions covering transcript generation, medical-record generation, and paperwork.

Tech, products & figures

aiDAPTIV+ (Phison's GPU memory-expansion solution)
Uses NAND Flash SSDs plus system DRAM as a GPU memory-expansion layer, breaking past VRAM capacity limits.
HBM vs. NAND Flash per-GB cost comparison
About US$200 vs. US$0.1.
LLaMA 3.2 (including 1B/3B/90B multimodal versions)
Used repeatedly throughout the talk as a cost example based on this model family's training requirements.
NVLink/400G high-speed networking vs. 10G ordinary networking
The network downgrade approach under Phison's architecture.

Notable quotes

"So is NVIDIA really selling compute, or selling RAM? They just want to sell both together."
"If you have an on-premises model, once it's trained... that is the ultimate right — so I encourage everyone to train their own models on-premises; you get to control everything."

Q&A

  • No Q&A record.

Fact-check notes

  • Verified via WebSearch: Phison's aiDAPTIV+ is a real product, and the official description matches the talk's content — 'an SSD designed with Phison's NAND expertise, controlled via aiDAPTIVlink, that splits and transfers data to the GPU for processing' — used to break past GPU memory-capacity limits. At the time of this 2024 conference talk, the product's name was indeed aiDAPTIV+; official press releases show that in 2026 Phison further upgraded the technology into the 'aiDAPTIV multi-tier memory architecture' and launched dedicated Pascari SSDs, extending to iGPU PC platforms — this is a later product iteration after this talk, not part of the content at the time of the 2024 talk, and the time gap is noted here for clarity.Sources:經濟日報群聯官網新聞稿TechNews
  • Statements such as 'SSDs sent to the space station and retrieved back on Earth' and 'cooperation with NASA on a lunar data center' are the speaker's own account, with no external corroboration found in independent news sources for the details — flagged as unverified in detail, though consistent in direction with Phison's long-standing publicity around aerospace-cooperation cases.

中文版