<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>Pradyun N.'s Writings</title>
<description>Some of my longer and shorter thoughts, all faithfully written in Org Mode.</description>
<generator>Emacs webfeeder.el</generator>
<link>https://pradyun.net</link>
<atom:link href="https://pradyun.net/rss.xml" rel="self" type="application/rss+xml"/>
<lastBuildDate>Thu, 30 Jul 2026 03:08:08 +0000</lastBuildDate>
<item>
  <title>Gallivanting Around With a New Camera</title>
  <description><![CDATA[<div id="content" class="content">

 <p>
Earlier this month I went camping in the  <a href="https://en.wikipedia.org/wiki/North_Cascades_National_Park">North Cascades</a> with my undergraduate roommates. Extremely, extremely happy with how that trip turned out – basically ~3 days of not worrying about research and resetting my brain, and it's been fairly positive for my productivity and mental health since I've gotten back.
</p>

 <p>
Importantly, we visited a camera store to get a tripod as part of the preparations for that trip. Lo and behold, I came across a used Fujifilm X-Pro 3 that cost an arm instead of an arm  <i>and</i> a leg. I'd been eyeing the body as an upgrade from my X-T20 for well over a year at this point, and this felt like the exact kind of trip I'd wanted the camera for. So caution went to the wind, and I walked out with a 15$ tripod and a new camera!
</p>

 <p>
Call it shiny object syndrome, but I've been extremely happy with the body since I've gotten it. I know that having to flip out the back screen is a source of annoyance for many people, but I've honestly enjoyed it? I had a phase in junior/senior year where I enjoyed shooting film on a Minolta XG-1, and really fell in love with the idea of "shooting and moving on" without the option of "chimping" shots through the back screen. 
</p>

 <p>
Sadly, the main drawbacks of film for me were the things most people  <i>love</i> about it (development time, inflexibility w.r.t ISO and white balance). The fact that you can't see the screen of the X-Pro 3 by default pretty strongly deters chimping, and while you can absolutely achieve the same effect with other cameras through discipline and willpower, it's definitely increased the immersion I feel when shooting. I do flip out the screen from time to time to play with film simulations and white balance, but I've been enjoying the practice of coming home with very little idea of how my shots turned out and going through it on my computer as a wind-down routine.
</p>

 <p>
The hybrid OVF/EVF rangefinder is also  <i>killer</i> – I can see why people pay so much for it on the X100VI! Whether I think the price is proportional to the value on that camera is a different question (I do not think so), but it's nice to be able to both preview what my shot will look like while having the option to see what "slice" of my FoV I'm capturing. I have a fairly bad habit of getting very sucked into the EVF and missing the forest for the trees, so swapping into the OVF when first composing a shot has helped. There's been more than a couple of times that I've realized that I want to clip something off that was at the fringe of the frame, or that there's something slightly out-of-frame that I'd prefer to re-compose with. 
</p>

 <p>
I do know that some people have issues with the rangefinder feeling a little odd ergonomically, but I've actually preferred it to the X-T20, as my arms feel a little less crammed into my chest. Though that is also probably partially due to the camera being wider. That does come with some additional heft (plus the fact that the construction is now titanium instead of plastic), but it hasn't been prohibitively heavy. I made it through   <a href="https://www.alltrails.com/trail/us/washington/maple-pass-trail">a fairly strenuous hike</a> (by my standards, at least) while carrying this camera and a spare lens, and have also been fairly comfortable walking around campus with it slung around my neck.
</p>

 <p>
So yeah, overall, very  <i>very</i> happy with this camera. I've been making more of an active effort to get out and go on photo-walks while campus is deserted (and social anxiety is slightly lower), and it's been an absolute joy to shoot on this thing. 
</p>

 <p>
For the record – do I think that I  <i>wouldn't</i> have enjoyed shooting on my old X-T20? Absolutely not. I do think that there are some shots where it would have been harder to shoot on it by virtue of the fact that I might have gotten too absorbed into the EVF, but that's not really an indictment of the camera. Photography in general is enjoyable more because of the creative process than the gear, and the X-T20 doesn't really limit that process in any meaningful way. Most modern cameras (post-2012, I want to say?) are capable of getting out shots with sufficient pixel density and color accuracy, so you might as well get whatever the cheapest camera is that you enjoy shooting with. But there's also something to be said for having a camera that makes you  <i>want</i> to get up and take pictures with it, and for a variety of reasons the X-Pro 3 evokes a fairly strong feeling of that from me.
</p>

 <p>
Anyway, I also wanted to share some highlights from the camera roll from our trip (the main reason this is categorized Long and not Short). Enjoy my show-and-tell into the internet void! Most shots below were taken on my beloved XF35mm F1.4, with some additional landscape shots taken on the Sigma 18-50mm F2.8. Everything is SOOC, so a crop-n-tilt and color adjustments may be due in the future.
</p>
 <div id="outline-container-org5a505d2" class="outline-2">
 <h2 id="org5a505d2">North Cascades Pictures</h2>
 <div class="outline-text-2" id="text-org5a505d2">
 <p>
I was the designated cameraman <sup>™</sup> for the roommates' Seattle trip (largely through self-imposition, for the record). We went camping for three days – the first was spent packing/road-tripping to the campsite, the second was spent hiking, and the third was spent kayaking. I mostly took pictures on the second day because I was too scared to take a new camera into the middle of a lake, and the pool of water in my kayak asserted that to be a good decision. I did get some fun shots on our way into the national park though.
</p>


 <div id="orga2e36b9" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79421-1920.webp" alt="XSB79421-1920.webp" class="halfwimg"></img></p>
</div>

 <p>
Amateur tip: sitting shotgun and setting shutter speed to 1/2000 or higher can get you some really nice pictures on scenic road trips, clean windshield permitting.
</p>


 <div id="orga2fd9e8" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79448-1920.webp" alt="XSB79448-1920.webp" class="halfwimg"></img></p>
</div>

 <p>
I was pretty shocked when we hit our first viewpoint on the drive in. Apparently water can turn this tealish-green color because of glacial sediment – the more you know!
</p>


 <div id="org7f40aac" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79455-1920.webp" alt="XSB79455-1920.webp" class="eightywimg"></img></p>
</div>

 <p>
This was also the closest I'd been to a snow-capped mountain in a long time, so I spent a solid amount of time staring up at those.
</p>

 <p>
The main event was a loop of the  <a href="https://www.alltrails.com/trail/us/washington/maple-pass-trail">Maple Pass trail</a>. The route is absolutely beautiful, albeit a bit steep for sedentary grad students (I wonder who that could be?). I started feeling a solid amount of struggle about a third in, but the first lookout was a much-needed shot in the arm.
</p>


 <div id="org2243ae7" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79490-1920.webp" alt="XSB79490-1920.webp" class="halfwimg"></img></p>
</div>

 <p>
This picture is somewhat emblematic of some of the struggles I had taking pictures throughout the hike. I wanted to get as much of the landscape as possible, which is inherently a very  <i>deep</i> thing to shoot, so I resorted to vertical shots a lot. At the same time, properly exposing the snow, greenery, and rock faces all at once was a tall task, and not one I'd read the manual enough at the time to tackle. So there are quite a few ethereally blown-out snowcaps and clouds.
</p>

 <p>
I did, however, get this shot of my friends taking a break which I thought was pretty cool. You can once again see that exposure was a challenge.
</p>


 <div id="orgf62eb06" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79492-1920.webp" alt="XSB79492-1920.webp" class="eightywimg"></img></p>
</div>

 <p>
It's not super clear from the pictures because I wasn't taking a lot of them initially and didn't want to hold up my friends, but the first part of our hike was through a really twisty forest section. At some point we broke out of the forest and entered a section I lovingly referred to as "Lord of the Rings" because of the way the greenery and lighting looked. At this point I also received reassurance that it was okay if I hung back a little and took my time taking pictures, so I got some decent shots of the terrain while we hiked.
</p>


 <div id="orgf680116" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79510-1920.webp" alt="XSB79510-1920.webp" class="eightywimg"></img></p>
</div>


 <div id="orgc1dc7df" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79520-1920.webp" alt="XSB79520-1920.webp" class="eightywimg"></img></p>
</div>

 <p>
A little farther into lord-of-the-rings, and we started to see some pretty huge rock faces.
</p>


 <div id="orgadf7698" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79538-1920.webp" alt="XSB79538-1920.webp" class="eightywimg"></img></p>
</div>

 <p>
We didn't realize it at the time, but we hit the best lookout/vantage point a little bit after that. It's a little hard to recollect exactly how I felt seeing this with my own two eyes, but quite a few "holy shits" and "oh my gods" were muttered while sitting on a rock.
</p>


 <div id="orga18c14c" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79593-1920.webp" alt="XSB79593-1920.webp" class="eightywimg"></img></p>
</div>


 <div id="org5c84476" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79598-1920.webp" alt="XSB79598-1920.webp" class="eightywimg"></img></p>
</div>

 <p>
This was also about the point that we realized that we still had about half the hike to go (a little more upwards then the entire downclimb), so we all picked up the pace a little at that point. As a result, the number of pictures taken declined rather rapidly, but there are a couple shots that I'm glad I took during the hurry.
</p>


 <div id="orgfc7f178" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79680-1920.webp" alt="XSB79680-1920.webp" class="halfwimg"></img></p>
</div>

 <p>
As an aside, we'd mostly been hiking through in tshirts and pants by this point. So staring at what seemed to be a scene straight out of an Indian movie's Europe shooting while wearing my standard work attire was pretty crazy.
</p>


 <div id="org485046d" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79692-1920.webp" alt="XSB79692-1920.webp" class="halfwimg"></img></p>
</div>

 <p>
This shot has a snakey leading line going through it that I liked. Plus my friend using his hiking stick as a golf club was kinda funny.
</p>

 <p>
And arguably my favorite shot from this trip:
</p>


 <div id="org60bf64b" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79695-1920.webp" alt="XSB79695-1920.webp" class="eightywimg"></img></p>
</div>

 <p>
In hindsight I might have cropped off a little more of the sky, but it gives the people some breathing room, which I kind of like. Coincidentally this was also the  <i>last</i> decent shot I took of our hike, since the footing was quite unstable during the downclimb and we were  <i>really</i> trying to beat the light at that point. Case in point, here's a weirdly exposed shot of the sunset light before we'd even made it back into the forest.
</p>


 <div id="orgef8c1e1" class="figure">
 <p> <img src="https://pub-b14bb3b617884985afa3a39f74640d0f.r2.dev/images/XSB79710-1920.webp" alt="XSB79710-1920.webp" class="halfwimg"></img></p>
</div>
</div>
</div>
 <div id="outline-container-org0fe3406" class="outline-2">
 <h2 id="org0fe3406">Closing Thoughts</h2>
 <div class="outline-text-2" id="text-org0fe3406">
 <p>
If it wasn't already apparent, we made it back before dark, and no major issues were had getting back to our campsite. The kayaking trip the next day was just as enjoyable, if a little less picturesque, though there are fewer visual records of that excursion.
</p>

 <p>
If I were to summarize this post so far:
</p>

 <blockquote>
 <p>
 <b>TL;DR:</b> X-Pro 3 good, did not know how to use it fully during the trip so pictures were less than ideal, North Cascades are absolutely beautiful.
</p>
</blockquote>

 <p>
I'd actually bought the camera right before we started driving, so I didn't have a ton of time to figure out the features of it that were different than my X-T20 and generally didn't have a great feel for how to best take advantage of it. The weekend after I got home, I finally sat down and read through the entire manual, configuring all the settings in a way that felt intuitive to me. This helped feel more comfortable when I went out on photowalks, and I've been taking pictures around campus that I'm much more satisfied with from a quality perspective.
</p>

 <p>
This post was initially intended to be a super-cut of all the pictures I'd enjoyed taking this month, including the improved ones from when I got back to campus, but the hike section ended up being way longer than I'd initially anticipated! Plus it's kinda fun recounting the experiences and thoughts that came along with the shots.
</p>

 <p>
I'm hoping to make a second post sometime in the next few weeks about the pictures I've taken since coming back to campus, since I had a bit more creative freedom when taking those to play with framing and whatnot.
</p>

 <p>
In the meantime, I hope this post was some mixture of informative and entertaining. It's a beautiful world out there, and we ought to take more pictures of it!
</p>
</div>
</div>
</div>]]></description>
  <link>https://pradyun.net/blog/new_camera.html</link>
  <guid isPermaLink="false">https://pradyun.net/blog/new_camera.html</guid>
  <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
</item>
<item>
  <title>Studying Linux Schedulers, and Why Metrics Matter</title>
  <description><![CDATA[<div id="content" class="content">

 <p>
Last semester, I took  <a href="https://courses.grainger.illinois.edu/cs533/sp2026/">Parallel Computer Architectures</a> at UIUC. For a term project, I (with the help of  <a href="https://screamingpigeon.net/">smart</a>  <a href="https://atmospheal.com/">friends</a>) set out to investigate exactly how much scheduling decisions mattered in the Linux Scheduler, specifically as it related to picking between different cores available in a CPU (or multiple sockets). 
</p>
 <div id="outline-container-org0762721" class="outline-2">
 <h2 id="org0762721">Background</h2>
 <div class="outline-text-2" id="text-org0762721">
 <p>
The classical advice goes that you prefer to keep processes on the same core for as long as possible (without inducing unfairness) so that you can keep their data warm in the caches, and Linux's EEVDF scheduler does a pretty good job of that. What  <i>isn't</i> taken into account is making sure processes are arranged across the CPU(s) in such a way that you minimize the number of levels in the cache hierarchy you have to traverse to service accesses to shared memory.
</p>

 <p>
As an illustrative example – consider a dual-socket Xeon system (which was our testing configuration). If a process makes an access to shared memory, let's be pessimistic and suppose that  <i>another</i> process has recently accessed this memory as well, has performed a write, and that the dirty data is still in the cache hierarchy. What are the possible coherence actions that could occur? It depends on where the process was scheduled:
</p>

 <ol class="org-ol"> <li>If the process was running on the  <i>same</i> core (in a separate scheduling quantum), nothing! We have the data! Yippee!  We will find it somewhere in our L1/L2 private cache hierarchy, and require no coherence actions.</li>
 <li>If the process ran on  <i>another</i> core, then on modern Xeon systems there are two ways this can pan out.
 <ol class="org-ol"> <li>The process was scheduled onto another core in  <i>our</i> socket. This means that we will force a write-back from the other core's private cache hierarchy into the fully-shared L3, and either pull data from DRAM or directly from the other core's private cache hierarchy via the coherency fabric.</li>
 <li>The process was scheduled on the  <i>other</i> socket. This is the worst-case, as the other socket's L3 will need to transfer data across the chipset (or we go to DRAM), at which point we can pull the data all the way through our own cache hierarchy to use it… yikes!</li>
</ol></li>
</ol> <p>
As we can see, there is a fairly asymmetric penalty here to service a cache miss when dealing with multi-threaded programs. So really, our goal is to minimize the number of times we have to deal with case 2.2. If we were only read-sharing across cores, everything would be hunky-dory since data is much more likely to stay cache-resident. It's really this pessimistic write-sharing behavior that starts to be annoying. Ideally, we generally prefer to trigger case 2.1 in this situation, since it balances cache miss penalties with scheduling flexibility and maximizing available compute. 
</p>

 <p>
The granularity of this decision tree will inevitably change depending on the CPU/configuration we are using. Zen 1 and 2, for example, would have a third branch here, as some SKUs contained separated L3 caches per CCX, so you could get away with faster access times if you managed to schedule your data-sharing processes on the same CCX. 
</p>

 <p>
So with some fairly strict caveats – no sub-NUMA clustering, ignoring memory residency, etc. – we set out to answer three questions:
</p>
 <ol class="org-ol"> <li>How much of a benefit is there from successfully scheduling multi-threaded processes fully within one LLC domain?</li>
 <li>Do current schedulers do a good job of ensuring that this happens?</li>
 <li>How can they be improved to actually make this happen?</li>
</ol> <p>
Our study mostly focused on points 1 and 2, with 3 being a "we hope this can happen", until we realized that  <a href="https://lore.kernel.org/lkml/cover.1775065312.git.tim.c.chen@linux.intel.com/">CAS</a> was starting to address point 3. So at this point we thought, "what if we just test CAS and EEVDF and find the gaps?". Okay, fair enough. But then we realized that we couldn't actually do that on shared lab resources and instead pivoted to profiling  <a href="https://lpc.events/event/19/contributions/2099/attachments/1875/4020/lpc-2025-lavd-meta.pdf">LAVD</a>. It obviously exhibits very different behaviors and priorities, but the hope was that the  <i>second-order</i> effects of LAVD would emulate that of CAS. I'm not totally sure how accurate that hypothesis ended up being, but EOTD we needed  <i>some</i> testing methodology so this would have to do.
</p>

 <p>
There's a lot more background about the internals of these schedulers about how various Linux schedulers work and our evaluation methods, but I'll fast forward to the mishaps in the next section for (relative) brevity. More fine-grained details can be found in our final report, either in  <a href="../static/scheduler_profiling.pdf">PDF</a> or  <a href="./scheduler_profiling.html">web</a> format.
</p>
</div>
</div>
 <div id="outline-container-orgd7d3d8f" class="outline-2">
 <h2 id="orgd7d3d8f">A Cascade of Mistakes</h2>
 <div class="outline-text-2" id="text-orgd7d3d8f">
 <p>
So at this point, our study is clear! We want to characterize how performance is impacted by scheduling threads auto-magically with the default Linux scheduler, with a "cache-optimized" scheduler, and with manual restriction of thread positions. 
</p>

 <p>
As many school projects go, we started a little close to the deadline, and followed a fairly linear train of thought to start testing right away. Our experimental setup was to track coherence actions (via read-for-ownership requests sent across the socket) and to track latency figures with the various configurations, and try to drill down exactly what aspect of runtime changing had contributed to performance increases/decreases. Since our benchmarks were all run with a thread count at/under the capacity of a single socket, our proxy for an "idealized" placement across cores was to simply restrict benchmarks to a single socket via  <code>numactl</code> as a control.
</p>

 <p>
This led to one of the most striking results of our entire study: both the standard Linux scheduler and specialized LAVD scheduler perform approximately the same on our benchmarks, but using  <code>numactl</code> to restrict thread scheduling to a single socket improved runtime by as much as 3x!
</p>


 <div id="orgcdc74c3" class="figure">
 <p> <img src="../images/scheduler/runtime_speedup_vs_eevdf_2s.png" alt="runtime_speedup_vs_eevdf_2s.png" width="80%"></img></p>
 <p> <span class="figure-number">Figure 1: </span>Normalized Speedup of Benchmark Runtime</p>
</div>

 <p>
So I guess if any of you are currently using a multi-socketed machine to run workloads in an under-subscribed fashion, I would recommend at least trying to use  <code>numactl</code> to pin workloads to a single LLC domain or a set of cores – you can gain  <i>extreme</i> speedup in multi-threaded shared-memory programs. This is mostly due to the fact that the Linux load balancer will try to ensure that all sockets are approximately equally loaded as thread count permits, even if that actively disadvantages the host program.
</p>

 <p>
The salient exception to this observation was the  <code>mediawiki</code> benchmark, which we sourced from DCPerf – how had the runtime barely moved at all? Maybe it was the fact that cache hitrates hadn't really improved? 
</p>

 <p>
Okay, we can check whether that's true. Let's take a look at if the number of read-for-ownership misses (e.g L3 wants to write to something but it doesn't have the data, so it asks the other core to let it own data) has changed. When we have strong write sharing, we expect that number to go up since shared cachelines should bounce between the two sockets' L3 caches.
</p>


 <div id="org7c3a8df" class="figure">
 <p> <img src="../images/scheduler/l3_rfo_mpki_newdata.png" alt="l3_rfo_mpki_newdata.png" width="80%"></img></p>
 <p> <span class="figure-number">Figure 2: </span>L3 Cache Read-for-Ownership Misses (per Kilo-Instruction)</p>
</div>

 <p>
Whelp. Now we have  <i>two</i> mysteries.
</p>

 <ol class="org-ol"> <li>How has  <code>mediawiki</code> not sped up at all, even with an improvement in L3 RFOs?</li>
 <li>What the hell is going on with  <code>perlbench</code>? What is making it speed up so much?</li>
</ol> <p>
We didn't end up being able to answer either of these questions prior to the deadline unfortunately – in fact, we found this relationship the day our paper was due! So with a severe sleep deficit and copious amounts of caffeine in our bloodstreams, we wrote that we hadn't been able to pin down the issue in our final report, and continued slogging through more graphs and writing. We ended up finishing the report at exactly 11:45PM, scampered over to Siebel and slid the report under our professor's door as fast as possible. And I don't think any of us really gave it another thought after submission.
</p>

 <p>
Cut to a couple weeks after end-of-semester, and I was cleaning up the traces we'd used for this project. And for some reason I have since forgotten, I was curious what the MIPS was for each of these benchmarks.
</p>


 <div id="orgd0bd4c7" class="figure">
 <p> <img src="../images/scheduler/mips_newdata.png" alt="mips_newdata.png" width="80%"></img></p>
 <p> <span class="figure-number">Figure 3: </span>Throughput Measured as Millions-of-Instructions per Second (MIPS)</p>
</div>

 <p>
Well that makes… no sense. Throughput has gone up significantly but runtime hasn't changed?? What command did we use to run the benchmark?
</p>

 <div class="org-src-container">
 <pre class="src src-bash"> <span style="font-weight: bold;">cd</span> /home/pradyun/pg_work/dcperf/
sudo ./perf_collect_mw  <span style="font-style: italic;">"$JOBNAME"</span> ./benchpress_cli.py run oss_performance_mediawiki_mlp  <span style="font-style: italic;">\</span>
    -i  <span style="font-style: italic;">'{"scale_out": 1, "client_threads": 32, "duration": "10m"}'</span> &
</pre>
</div>

 <p>
… Oh boy. Duration was fixed to 10 minutes. This directive was buried so deeply in our testing infrastructure, and reused  <i>so</i> many times over the course of this project that we'd completely forgotten about it.
</p>

 <p>
So in one shot, several mistakes have been made. Speedup was calculated with the assumption that instructions have been held constant, which was not the case for this benchmark. And not only have we accidentally reported another single-socket speedup as having been a phantom "no-change", we've also somehow  <i>masked the one datapoint that made LAVD look better</i>. For the latter bit, my theory is that the nature of  <code>mediawiki</code> suits the chain-of-tasks model that LAVD was designed for.
</p>

 <blockquote>
 <p>
 <i>Side Note</i>: LAVD was designed and optimized for the Steam Deck by Changwoo Min and the great folks over at  <a href="https://www.igalia.com/">Igalia</a>. He had a pretty great talk about LAVD at OSPM25 which you can find  <a href="https://www.youtube.com/watch?v=pHX-KgBCjwQ">here</a>, would recommend watching!
</p>
</blockquote>
</div>
</div>
 <div id="outline-container-org1c4193f" class="outline-2">
 <h2 id="org1c4193f">Conclusions</h2>
 <div class="outline-text-2" id="text-org1c4193f">
 <p>
So that was… long-winded. I hope it was at the very least entertaining! Regardless, this is one of those experiences that I hope sticks with me. For one, I hope it teaches me to stop procrastinating schoolwork so much sometimes, but also that the choice of metrics can very easily disguise results as something they're not. That can both be in the negative direction (this study) but also the positive direction. Both of these directions can be pretty dangerous, especially in environments where you fix the benchmarks and tests early on and forget about them going forwards (e.g performance engineering CI pipelines). I guess the long and short of what I'm trying to say is this:
</p>

 <blockquote>
 <p>
Picking good metrics to communicate your results helps not only present findings better, but also lets your work be resilient to mistakes or oversights made in experimental design.
</p>
</blockquote>

 <p>
At a more meta-level, I feel kind of obligated to take more care with experimental methodology in future profiling studies. This was really the first time I'd profiled anything on a computer without pre-rolled tooling to do it (e.g NVIDIA's SoL tools), and it definitely showed in some of the rougher edges of our evaluation. But at the very least, I'm glad that this saga occurred within the confines of a school project rather than an actual research paper, and that I can carry the lessons here into future research.
</p>

 <p>
If you're interested in some of the finer-grain details of our evaluation methodology, more performance counters from our testing, or a fairly long expository section about how LAVD and EEVDF work in Linux, I'd recommend reading our  <a href="./scheduler_profiling.html">final report</a>. It's not the most well-refined study ever, but we all learned a lot from it!
</p>

 <blockquote>
 <p>
 <b>AI Disclosure</b>: This post was written by hand and then later proofread/fact-checked with the help of an LLM. I have been an abuser of  <code>--</code> for longer than I can remember, and I will continue to use it – unfortunately, it gets converted into an em-dash by my Org export engine! It's truly unfortunate that hyphenation has been broadly conflated with "AI slop", yet I will continue to use it at will.
</p>
</blockquote>
</div>
</div>
</div>]]></description>
  <link>https://pradyun.net/blog/metrics_matter.html</link>
  <guid isPermaLink="false">https://pradyun.net/blog/metrics_matter.html</guid>
  <pubDate>Sat, 18 Jul 2026 00:00:00 +0000</pubDate>
</item>
<item>
  <title>Use It or Lose It</title>
  <description><![CDATA[<div id="content" class="content">

 <p>
In Spring 2024, I took  <a href="https://courses.grainger.illinois.edu/ece513/sp2024/index.html">ECE 513</a>. To this day, I consider it the hardest class I've ever taken, and also the most rewarding in many ways. I'd always wanted to take a hard math class, and somewhat romanticized the notion of agonizing over  a difficult problem set. Rest assured, I no longer romanticize math classes as such, and am indebted to the (very) smart classmates that endured  <i>many</i> nights of me confusing myself. Though in retrospect, I did quite enjoy the puzzle nature of solving difficult problems, and hope to take another math class in grad school (most likely some form of abstract algebra).
</p>

 <p>
Despite the trials of ECE 513, I seem to have utterly forgotten most of the formulas and theorems we learned in the class a mere 2 years later! Reading my old homeworks, I can  <i>vaguely</i> understand the language and concepts at play, but if you asked me to derive the solutions again, I probably wouldn't have the foggiest idea how to do so! Some high-level concepts and a general familiarity with linear algebra have stuck with me though, and those are arguably  <i>very</i> nice skills to have in the AI era.
</p>

 <p>
Obviously abstract linear algebra and DSP are not my fields of research/work, but I apparently also forgot the chain rule at some point? Which I found out while trying to learn about how backpropagation works. And calculus fundamentals feel like a fairly important thing to retain for general knowledge purposes.
</p>

 <p>
This note isn't really supposed to have any grand meaning, but some thoughts:
</p>

 <ol class="org-ol"> <li>I'm (now even more) incredibly awed by the professors and grad students in my life who seem to have an almost eidetic recall of the coursework and papers they've studied in the past.</li>
 <li>What is the "correct" amount of effort to spend reviewing or reinforcing content learned in past courses without impeding current work or research?</li>
 <li>I wonder if limited recall is somehow attributable to the proliferation of search engines and online databases? Something something "incentivizing  <i>finding</i> information repeatedly instead of recalling information long-term".</li>
</ol> <p>
Can't imagine LLMs are doing anything to help point 3… 
</p>
</div>]]></description>
  <link>https://pradyun.net/notes/use_it.html</link>
  <guid isPermaLink="false">https://pradyun.net/notes/use_it.html</guid>
  <pubDate>Fri, 19 Jun 2026 00:00:00 +0000</pubDate>
</item>
<item>
  <title>A Brief Review of the Kobo Libra Color</title>
  <description><![CDATA[<div id="content" class="content">

 <blockquote>
 <p>
TL;DR I'm a fan! It's light, it's got a nice screen, it's got good battery life, and it reads EPUBs natively.
</p>
</blockquote>

 <p>
More details below, in case the hyper-descriptive summary above wasn't enough for you.
</p>

 <p>
Screen size/resolution hasn't been a huge issue for me, though that may be because I'm fairly partial to small font sizes. Kobo's operating system also natively supports uploading and using custom fonts, which is rather kind of them – I've opted to use Crimson Font, as I usually do for serif typefaces. I've read through some longer books ( <a href="https://www.kobo.com/us/en/ebook/how-to-build-a-car-the-autobiography-of-the-world-s-greatest-formula-1-designer?sId=efc0a036-89ff-4e9f-95c8-219b60b8c4cc">How to Build a Car</a>,  <a href="https://www.kobo.com/ww/en/ebook/the-song-of-achilles">The Song of Achilles</a>) without issue, and it's been fairly enjoyable to read with my Kobo in one hand and a cup of coffee in the other.
</p>

 <p>
The buttons on the side are likely a design decision lifted from the Kindle Oasis, and I'm totally okay with that. I had my eyes on the Kindle Oasis for years but never thought I was reading ebooks enough to justify the upgrade from my Paperwhite. Sadly Amazon has gone in a  <i>different</i> direction these days so the Oasis was discontinued. Regardless, the Libra Color is fairly similar in spirit, and has the benefit of using a color panel at a cheaper price than what Amazon charges for the Colorsoft. Some semi-evangelical forum members (cough cough, Reddit) claim that color panels can lead to meaningful loss in DPI and screen clarity, but I have noted no such issues personally speaking.
</p>

 <p>
There's also been some active communities that have worked on modding and adding extensions to the Kobo Libra – NickelMenu being a prolific example – but I haven't actually found significant utility from any of them. I was hoping to find a way to read PDFs comfortably on the Libra, but that hasn't been very fruitful given the smaller screen size and fixed-layout properties of PDF. There doesn't seem to be any (recent) software that can convert PDFs into EPUB fairly accurately, so in the meantime I've been sticking to using Zotero on my desktop or iPad for paper-reading. This is also probably the more defensible way to do it long-term, purely from a perspective of managing annotations. 
</p>

 <p>
Would I buy it again? Definitely. Should you buy it? Debatable. Will you regret it? Assuming you actually read on it, I'd find it unlikely.
</p>
</div>]]></description>
  <link>https://pradyun.net/notes/kobo.html</link>
  <guid isPermaLink="false">https://pradyun.net/notes/kobo.html</guid>
  <pubDate>Sun, 31 May 2026 00:00:00 +0000</pubDate>
</item>
<item>
  <title>Installing Vivado on NixOS</title>
  <description><![CDATA[<div id="content" class="content">

 <p>
This note started out in May 2025, when I first tried to install Vivado on my NixOS machine. I'd tried all sorts of random things. Below you will see some bullet points from the  <i>original</i> draft of this post, which was intended to be much longer, as a triumphant victory after wrangling the mess of scripts and dependencies that is Xilinx packaging.
</p>

 <hr></hr> <ul class="org-ul"> <li>Dump installation image with web installer (SFD not supported post-2025.2)
 <ul class="org-ul"> <li>Make an FHS to launch web installer, scripts are broken for  <code>nix-ld</code></li>
</ul></li>
 <li>Add dependencies to nix-ld</li>
 <li>Fix shebangs in scripts in  <code>bin</code></li>
</ul> <hr></hr> <p>
After about 4 days of agony I stumbled upon  <a href="https://gitlab.com/doronbehar/nix-xilinx">this repo</a> for porting Xilinx to Nix… even the author had given up and switched to Ubuntu! By his recommendation I threw in the towel and created a  <a href="https://gitlab.com/pradyun/nix/-/blob/master/containers/vivadobox.ini?ref_type=heads">distrobox container</a> for my Vivado install, which has worked flawlessly, carrying me through synthesis of fairly large designs for the VCU118 during my chip tapeout. I'll write about that experience in the future (I hope), but in the meantime the moral of the story is:
</p>

 <blockquote>
 <p>
NixOS is pretty cool, until you use EDA software, at which point you might as well just use Ubuntu.
</p>
</blockquote>
</div>]]></description>
  <link>https://pradyun.net/notes/vivado_nixos.html</link>
  <guid isPermaLink="false">https://pradyun.net/notes/vivado_nixos.html</guid>
  <pubDate>Sat, 18 Apr 2026 00:00:00 +0000</pubDate>
</item>
<item>
  <title>A Primer on Oblivious RAM</title>
  <description><![CDATA[<div id="content" class="content">

 <div class="abstract" id="orgeb02040">
 <p>
We discuss the motivations and principles underpinning Oblivious RAM (ORAM) and the naive construction of one. We then transition to a more in-depth review of the drawbacks of older ORAM constructions, and how modern implementations like Path ORAM and Ring ORAM address them. Finally, we close on applications of these more efficient implementations, and their relevance to modern research in computer science and cryptography.
</p>

</div>

 <blockquote>
 <p>
 <b>Note</b>: This was a paper I wrote for Cryptography (ECE407) over the span of about 3 days straight. The literary quality dips in places and it's not my greatest work, but the topic overall is quite interesting. Hopefully you learn something new! 
</p>
</blockquote>


 <div id="outline-container-org712cf75" class="outline-2">
 <h2 id="org712cf75">Revisions</h2>
 <div class="outline-text-2" id="text-org712cf75">
 <ul class="org-ul"> <li> <i>05/22</i>: Client-side space overhead for Square Root ORAM was incorrectly characterized as \(O(N\log N)\)</li>
</ul></div>
</div>
 <div id="outline-container-org26e9aa6" class="outline-2">
 <h2 id="org26e9aa6">Introduction</h2>
 <div class="outline-text-2" id="text-org26e9aa6">
 <p>
Oblivious computation research dates back as far as 1979. Pippenger et al.  <a href="#citeproc_bib_item_1">[1]</a> explored the idea of an "oblivious Turing Machine", in which the motion of a Turing Machine's head on its tape lent no information about the underlying program. That is to say, you could make no useful conclusions regarding the input program by observing the head position - its movement was  <i>indistinguishable</i> for any two input programs of similar length. This paper not only laid out the foundations for a functional OTM, but it also was the first paper to discuss the notion of  <i>oblivious computing</i>. Computation in which the observable accesses and movements of a machine lent no information about the underlying operations.
</p>

 <p>
In their seminal 1987 papers (later reunited as a journal paper in  <a href="#citeproc_bib_item_2">[2]</a>), Goldreich and Ostrovsky extended the idea of oblivious computing to an off-CPU RAM that stored some sort of software. However, their motivation had more to do with software piracy. Since programs on a RAM were still just a sequence of bits, they could easily be copied and distributed by an adversary.
</p>

 <p>
The knee-jerk solution to this problem would likely be to simply encrypt the data going in and out of the RAM. However, it's important to note that  <i>where</i> and  <i>how</i> memory is accessed can lend important information about the underlying program, especially when paired with a high-level description of what the program does (which you get by buying the software). For example - repeated reads of a cluster of adjacent RAM blocks could indicate some sort of loop structure within the program itself. This is obviously subject to some architectural decisions on the CPU's memory subsystem, but the point should stand as a motivating example. Thus, a system allowing for oblivious memory access should be able to mask not only  <i>what</i> the data is, but  <i>where</i> the data is and  <i>how</i> the data was accessed.
</p>

 <p>
Today's compute landscape presents other scenarios requiring security when accessing a RAM, outside of the initial goal of piracy-prevention. A large motivating factor, as outlined in  <a href="#citeproc_bib_item_3">[3]</a> by Stefanov et al., is the advent of  <i>cloud compute</i>. Many clients store secure data (tax forms, legal documents, etc.) on remote servers that are considered "untrusted". We also heavily rely on cloud compute for many modern services, many of which will inevitably require some amount of confidential information from the user.
</p>

 <p>
In a naive security model where the client and cloud use symmetric encryption to prevent data from being leaked to eavesdroppers, the access pattern can still lend information regarding the type of actions that the client is engaging in with remote storage. Furthermore, the  <i>server</i> has knowledge of the physical memory locations that the client accesses, which inherently lends insights into the off-server actions of the client. Thus, in the early 2010s there was a renewed interest in devising secure memory schemes enabling oblivious compute with untrusted remote servers.
</p>

 <p>
 <a href="#citeproc_bib_item_2">[2]</a> was seminal for many reasons. Today, we recognize it as the paper that  <i>defined</i> the security model of  <b>Oblivious RAM</b>, and outlined the basic schemes of implementing it that motivated future ORAM schemes for nearly 2 decades.
</p>
</div>
 <div id="outline-container-orgefac073" class="outline-3">
 <h3 id="orgefac073">Problem Definition</h3>
 <div class="outline-text-3" id="text-orgefac073">
 <p>
The ORAM problem was first presented as two separate definitions in  <a href="#citeproc_bib_item_2">[2]</a>, found in sections 2.3.1 and 2.3.2. These refer to the definition of an ORAM to guarantee access pattern indistinguishability, then outline the definition of an ORAM "simulating" RAM as a way to prove correctness of the ORAM scheme. The aforementioned definitions are relatively obfuscated for the purposes of this primer (entrenched in Turing Machine models and whatnot), but there are some important takeaways to note.
</p>

 <ol class="org-ol"> <li>An ORAM system consists of a  <i>traditional</i> RAM block and an ORAM protocol being run on the attached CPU. The RAM itself doesn't do any of the "obliviation", so to speak. This makes ORAM a bit of a misnomer, as it describes the  <i>perceived function</i> of the RAM, not necessarily the function of the RAM itself.</li>
 <li>In any ORAM system, the client or CPU is considered  <i>secure</i>. Memories inside of the client (registers, caches) and internal operations of the client cannot be determined by an adversary – they can only see the traffic between the client and the memory, and the data stored directly inside of the memory.</li>
 <li>To avoid leaking information through deterministic encryption or decision-making over multiple runs of the same program, any CPU involved in a feasible ORAM system must include some sort of  random function and be able to make decisions based on it. In practice, this can be implemented as a PRF or PRG, with a randomly generated key or nonce used to vary the encryption of identical plaintext over different iterations.</li>
</ol> <p>
Modern papers like  <a href="#citeproc_bib_item_3">[3]</a>,  <a href="#citeproc_bib_item_4">[4]</a> prefer a simpler definition of the ORAM problem, and this is the one that we use going forwards.
</p>

 <blockquote>
 <p>
Consider a sequence of memory accesses \(\vec{y}\), defined below, where \(| \vec{y} | = M\).
</p>

 <p>
\[
\vec{y} = ((op_M, addr_M, data_M), ..., (op_1, addr_1, data_1))
\]
</p>

 <p>
\(op_i\) denotes whether the i-th access is a read or write, \(addr_i\) denotes the address of the i-th memory access, and \(data_i\) is the data written into the RAM for a write access.
</p>

 <p>
Given an ORAM algorithm, a sequence of logical client-server or CPU-RAM accesses \(\vec{y}\) is transformed into an observed access sequence \(\mathsf{ORAM}(\vec{y})\). ORAM protocols will guarantee that for any two access sequences \(\vec{y}, \vec{z}\) where \(|\vec{y}| = |\vec{z}|\), \(\mathsf{ORAM}(\vec{y}) \stackrel{c}{\approx} \mathsf{ORAM}(\vec{z})\). Furthermore, the data returned by \(\mathsf{ORAM}(\vec{y})\) must, with overwhelming probability, be consistent with the data returned by \(\vec{y}\) without the presence of ORAM (e.g "correct").
</p>
</blockquote>

 <p>
Throughout the rest of this primer, we refer to data in units of a single memory access, which we call a  <b>block</b>. We also, for brevity, embrace a "client-memory" terminology to encompass all possible systems that may implement an ORAM scheme. While it is possible to have multiple clients share data access while employing an ORAM scheme (see  <a href="#org5f3c112">Applications</a>), a single-client model is adopted in this primer to simplify discussion of encryption and client-side state.
</p>
</div>
</div>
 <div id="outline-container-org0888c6d" class="outline-3">
 <h3 id="org0888c6d">Implied Constraints of an ORAM</h3>
 <div class="outline-text-3" id="text-org0888c6d">
 <p>
Given our earlier definition, we can start to determine some of the properties of what an ORAM should look like. The following section skims some of the high-level analyses in  <a href="#citeproc_bib_item_2">[2]</a>, and some personal understandings from reading various papers in the field. For starters, we can separate adversaries into two types.
</p>

 <ol class="org-ol"> <li> <i>Non-Tampering Adversaries</i>: Eavesdrops on the transactions happening between the client and the memory.</li>
 <li> <i>Tampering Adversaries</i>: Can actively write bad data into the ORAM, and do everything that a non-tampering adversary might.</li>
</ol> <p>
Combating the additional actions of the Tampering adversary can be done via some sort of digital signature generated by the client that is stored with the data block. Reading a block with a bad signature could trigger some sort of error correction protocol. As a result, as alluded to earlier, the  <i>eavesdropping</i> problem becomes the more difficult one to solve efficiently.
</p>

 <p>
The indistinguishability criterion set forth earlier necessitates homogeneity of memory access pattern regardless of whether it is a read or write. Inherently, this means that every memory access will  <i>always</i> have to read something out of the ORAM. Block writes would need to be batched and deferred to some later event, or you would need to always write back blocks (even on read operations). At a high level, regardless of the access type made, the actions being taken by the client should  <i>look</i> identical.
</p>

 <p>
There is also the problem of masking repeated accesses to the same data block – if the blocks are naively read/written to the same location between accesses, then it becomes easy to detect data reuse. As a result, any practical ORAM protocol will need to relocate or shuffle data blocks over execution.
</p>

 <p>
Finally, since an ORAM breaks the implicit mapping of a block address or tag being used as a RAM index, we must attach additional metadata to each block in the ORAM to identify it after decryption. This metadata can include the block tag (encrypted with the block to avoid leaking mappings), a nonce used to randomize the encryption function, and any other properties that the ORAM protocol in question requires for correctness. It is generally assumed that any metadata required or suggested by a protocol is a fixed constant negligible to the size of the data block.
</p>

 <p>
Inherently, each of the above properties requires access redundancy or additional computation over a naive RAM. The goal of the ORAM implementations presented in this primer is to maintain them while maintaining computational feasibility and ensuring a negligible failure probability.
</p>
</div>
</div>
</div>
 <div id="outline-container-org79fcf63" class="outline-2">
 <h2 id="org79fcf63">An Early ORAM Construction</h2>
 <div class="outline-text-2" id="text-org79fcf63">
 <p>
With the above objectives, one can think of some basic implementations of ORAM. An especially trivial one would be "brute force" – regardless of the memory access, you can read every block in the memory into CPU registers, optionally modify it, re-encrypt it with a new nonce, then write it back into the same location of the RAM. If a block needs to be reused later, it can be kept in the CPU for operation even after linear scan.
</p>

 <p>
This scheme isn't all bad. For one, it doesn't require any special state about block mappings on the CPU – every block is exactly where you'd expect it to be, it's just that the server doesn't know which one you actually meant to use. However, the linear scan requires \(O(N)\) memory accesses given an N-block memory, and requires \(O(N)\) bandwidth. Any practical system would rattle to a halt. This motivated early research into devising an ORAM algorithm that was  <i>bandwidth-efficient</i> (and, as a corollary, memory-access-efficient).
</p>

 <p>
As if it weren't enough to design a security definition and objective for Oblivious RAM, Goldreich and Ostrovsky outlined early optimized ORAM schemes in  <a href="#citeproc_bib_item_2">[2]</a>. We discuss their initial proposal of a  <i>Square Root ORAM</i> and its pitfalls.
</p>
</div>
 <div id="outline-container-org9532a5c" class="outline-3">
 <h3 id="org9532a5c">Square Root ORAM</h3>
 <div class="outline-text-3" id="text-org9532a5c">
 <p>
This construction simulates an N-block RAM by using \(N + 2\sqrt{N}\) real memory blocks – an overhead of \(O(\sqrt{N})\) extra memory capacity (and the source of its name). Instead of brute-forcing the entire memory to hide its accesses, it  <i>permuted</i> the order of its blocks and periodically reshuffled them to prevent information leakage through repeated access.
</p>


 <div id="org629bf96" class="figure">
 <p> <img src="../images/sqrt-oram.png" alt="sqrt-oram.png" height="100pt"></img></p>
 <p> <span class="figure-number">Figure 1: </span>Memory Separation in Square Root ORAM</p>
</div>

 <p>
As shown in figure  <a href="#org629bf96">1</a>, the memory used for this scheme separates into three sections – the real blocks reserved for the actual data in the RAM, a shelter that acts like a "working memory", and dummy blocks (assume these are just set to 0) that help mask accesses when working with the shelter. At initialization, the shelter blocks are also assumed to be populated as dummy blocks.
</p>

 <p>
Square Root ORAM works in "epochs" of \(\sqrt{N}\) memory accesses – over this duration, blocks are taken out of the real blocks and placed into the shelter. After an epoch, in the worst case, all of the blocks in the shelter will have been filled up by legitimate blocks with potentially updated data. The overall algorithm for a single epoch is as follows.
</p>

 <ol class="org-ol"> <li>Using the random function, map each logical and dummy block index onto a random tag via a random oracle \(\pi\). Tags are assumed to be on the range \([1, \frac{n^2}{\epsilon}]\), where \(\epsilon\) determines our collision probability.</li>
 <li>Make a single pass over all the permuted blocks to update their tag values. It is assumed that there is either an inverse mapping available from the previous epoch's tags, or the blocks are given encrypted metadata that identify their intended logical index.</li>
 <li>Implement Batcher's Sorting Network (a binary tournament sorting algorithm) to sort the permuted blocks by tag value.
 <ul class="org-ul"> <li>This network is used since its access pattern is identical irrespective of input data. We call such sorting algorithms an  <i>oblivious sort</i>.</li>
</ul></li>
 <li>Simulate \(\sqrt{N}\) memory accesses.</li>
 <li>Implement Batcher's Sorting Network over the  <i>entire</i> memory, by the key \((v, \sigma)\). \(v\) is the index into the logical RAM (\(\infty\) for dummy blocks) and \(\sigma\) is 0 for permuted blocks, 1 for shelter blocks.</li>
 <li>Make a reverse scan through the ORAM, and mark each duplicate block as a dummy (e.g \(\sigma = 0\) for this block, but we saw \(\sigma=1\) the previous iteration).</li>
</ol> <p>
The final sort is a bit ambiguous, but it works since we assume that the dummy blocks  <i>always</i> have a tag \(\infty\). After step 5, these dummy blocks all end up in the final \(\sqrt{N}\) blocks – therefore any block with a valid tag must have been placed into the permuted blocks, effectively evicting the shelter. We also assume that the shelter will have at maximum one copy of each block (this will be clearer after explaining the access protocol), so after our sort we will find exactly one or two copies of each real block (hence \(\sigma\) allowing for a strict ordering).
</p>

 <p>
Each access \((op_i, addr_i, data_i)\) of step 4 is performed as follows, where \(i \in [1, \sqrt{N}]\).
</p>

 <ol class="org-ol"> <li>Iterate over all \(\sqrt{N}\) blocks in the shelter.
 <ul class="org-ul"> <li>Read each block into the CPU.</li>
 <li>Update it to \(data_i\) if \(addr_i\) matches the block's metadata and \(op_i\) is a write.</li>
 <li>Re-encrypt the block and write back into the shelter.</li>
</ul></li>
 <li>Select which block to read out of the permuted blocks.
 <ul class="org-ul"> <li>If the block was found in the shelter, we target \(tag_i = \pi(N + i)\) (the i-th dummy block).</li>
 <li>If we could not find it, we look for the real block by targeting \(tag_i = \pi(addr_i)\).</li>
</ul></li>
 <li>Since we sorted the permuted blocks in step 3 of the epoch, we can perform an oblivious binary search to read out the block with \(tag_i\).
 <ul class="org-ul"> <li>The blocks are permuted every epoch, and each permuted block is accessed once per epoch.</li>
 <li>No information can be leaked about whether a dummy or real block was read, or reused.</li>
</ul></li>
 <li>If the block was not found in the shelter and \(op_i\) is a write, update block data to \(data_i\). We then re-encrypt the block and write it into real block \(N + \sqrt{N} + i\) – e.g the i-th block of the shelter.</li>
</ol> <p>
At a high level, we can outlined scheme fulfills some of the basic properties we desire from an ORAM. Both access types and block reuse cannot be leaked to an adversary, as every memory access uses an observably identical protocol. The reshuffling of permuted blocks further guarantees that we do not leak any information regarding data locality.
</p>

 <p>
Furthermore, we can also see an improvement from our initial brute force scheme. The old scheme would have incurred a bandwidth of \(O(N\sqrt{N})\) per epoch. By incurring an additional client-side space overhead and a memory overhead of \(O(\sqrt{N})\), we reduce the per-epoch bandwidth to \(O(N\log^2N)\), which is dominated by the accesses required for Batcher's Search Network in steps 3 and 5 of each epoch.
</p>
</div>
</div>
 <div id="outline-container-org920cb7a" class="outline-3">
 <h3 id="org920cb7a">Drawbacks</h3>
 <div class="outline-text-3" id="text-org920cb7a">
 <p>
Despite being described as a "naive" solution by  <a href="#citeproc_bib_item_2">[2]</a>, this scheme presents some of the issues that plagued many ORAM schemes, the foremost one being  <i>complexity</i>. It is difficult to intuitively reason about the correctness and operation of this scheme, and becomes even harder so when  <a href="#citeproc_bib_item_2">[2]</a> later adapts it to a hierarchical solution.
</p>

 <p>
There is also the non-trivial performance overhead - as explained earlier, this scheme requires significant storage overhead in the memory of an ORAM system, and the per-access latency (\(\approx \log^2N\)) does not scale well with memory sizes. Goldreich and Ostrovsky argue in section 6 of their paper that there must be a lower bound of \(\Omega(t\log{t})\)  <sup> <a id="fnr.ackshually" class="footref" href="#fn.ackshually" role="doc-backlink">1</a></sup> real memory accesses to simulate \(t\) logical memory accesses on an ORAM.
</p>
</div>
</div>
</div>
 <div id="outline-container-orgfce42ea" class="outline-2">
 <h2 id="orgfce42ea">Efficient Tree-Based ORAMs</h2>
 <div class="outline-text-2" id="text-orgfce42ea">
 <p>
Throughout the early 2000s, many ORAM schemes were devised to cut per-access bandwidth closer to this boundary. Efficient ORAM constructions like Pinkas-Reimann  <a href="#citeproc_bib_item_5">[5]</a> and Goodrich-Mitzenmacher  <a href="#citeproc_bib_item_6">[6]</a> still inherited from a hierarchical RAM scheme presented in section 5 of  <a href="#citeproc_bib_item_2">[2]</a> and relied on oblivious sorting mechanisms. This meant that while amortized per-access ORAM latency approached logarithmic/polylogarithmic functions of N, worst-case latency was still on the order of \(\Omega(N)\) in  <a href="#citeproc_bib_item_6">[6]</a> or \(\Omega(N\log N)\) in  <a href="#citeproc_bib_item_5">[5]</a>. Altogether, this amounted in a performance overhead that ranged from 1400 to over  <i>80000</i> times the number of operations required for the same access pattern with a transparent memory device  <a href="#citeproc_bib_item_3">[3]</a> .
</p>

 <p>
While some ORAM schemes around this time, like the SSS construction in  <a href="#citeproc_bib_item_3">[3]</a> and Shi et al.'s \(\log^3\) ORAM  <a href="#citeproc_bib_item_7">[7]</a>, presented schemes that helped reduce the worst-case performance to polylogarithmic functions and reduce client/storage cost in doing so, tree-based constructions are currently the most popular category of ORAM. This is largely due to the innovations that came with the Path ORAM construction by Stefanov et al.  <a href="#citeproc_bib_item_8">[8]</a>.
</p>
</div>
 <div id="outline-container-org78ad35b" class="outline-3">
 <h3 id="org78ad35b">Path ORAM</h3>
 <div class="outline-text-3" id="text-org78ad35b">
 <p>
The general Path ORAM construction demonstrated in  <a href="#citeproc_bib_item_8">[8]</a> inherits heavily from the previous work done by Shi et al. in  <a href="#citeproc_bib_item_7">[7]</a>, particularly in its usage of bucket-based storage trees with oblivious eviction.
</p>


 <div id="org668644d" class="figure">
 <p> <img src="../images/path-oram.png" alt="path-oram.png" height="250pt"></img></p>
 <p> <span class="figure-number">Figure 2: </span>Path ORAM Memory Structure for \(Z=4, L=2\)</p>
</div>


 <p>
On the server side, storage is interpreted as an N-ary tree (typically  <i>binary</i>) of \(L\) levels, where each node is a 'bucket' storing up to \(Z\) real memory blocks. Buckets with less than \(Z\) blocks will be padded with dummy blocks. An example of such a tree is laid out in figure  <a href="#org668644d">2</a>. In the memory, these buckets would be laid out serially so they can still be accessed via a normal external indexing scheme. 
</p>

 <p>
Path ORAM, like most high-performance ORAM schemes, is also stateful on the client side. The client side must maintain a  <b>position map</b> (\(\pi\)) to help it find a desired block inside the memory, as well as a  <b>stash</b> (\(S\)) to store blocks in the client during runtime (more on this later). The position map is constructed such that indexing it with a block ID will return a leaf node in the memory's tree. The Path ORAM algorithm guarantees that when it traverses the tree to reach this leaf node, the desired block will be found in one of the buckets along this path (hence the algorithm's name).
</p>

 <p>
It is assumed that at initialization, the client's stash is empty, the server only contains dummy blocks, and the position map is filled with uniformly random numbers corresponding to various leaf node indices. With this in mind, we outline the Path ORAM algorithm below for a single memory access \(\vec{y}_i\). We adopt the notation \(P(x)\) to denote the path from the root node to leaf node \(x\), and \(P(x, i)\) to denote the i-th node on this path.
</p>

 <ol class="org-ol"> <li>Load \(\pi(addr_i)\) into \(leaf_i\), then reassign \(\pi(addr_i)\) to a random leaf node.</li>
 <li>Load each of the buckets along \(P(leaf_i)\).</li>
 <li>Add the desired block \(addr_i\) to the stash \(S\).</li>
 <li>If \(op_i\) is a write, then update block \(addr_i\) to contain \(data_i\) in \(S\).</li>
 <li>Iterate backwards along the path that was traversed (leaf node upwards, e.g \(P(leaf_i, L)\) to \(P(leaf_i, 0)\)).
 <ul class="org-ul"> <li>At each bucket, select up to \(Z\) blocks from the stash that have this bucket  <i>somewhere</i> on their path (per \(P(\pi(a))\) for some block \(a\)). Write them into the bucket.</li>
 <li>If we cannot find \(Z\) suitable blocks, write dummy blocks until the bucket is at capacity.</li>
 <li>Remove any blocks written to the bucket from our stash.</li>
</ul></li>
</ol> <p>
Like the Square Root ORAM  <a href="#citeproc_bib_item_2">[2]</a> and our initial brute-force solution, every access using this algorithm looks identical for any observer. \(Z(L+1)\) blocks are read out of the memory and \(Z(L+1)\) blocks are written into it along some randomly selected path. Since the blocks are, in effect, evicted from the tree when read out and later relocated to a random location, we incur no risk of leaking data locality.
</p>

 <p>
An interesting feature of this algorithm is the approach it takes to filling in the bucket tree at the end of each access. By taking a bottom-up traversion and pushing blocks as far down the tree as possible, buckets with fewer potential candidates are populated first. This in turn minimizes stash usage.
</p>
</div>
 <div id="outline-container-org9b86754" class="outline-4">
 <h4 id="org9b86754">Performance and Usability Implications</h4>
 <div class="outline-text-4" id="text-org9b86754">
 <p>
As every access for Path ORAM is algorithmically identical, there is no need to calculate amortized latency. Since \(L\) is set to \(\log_2 N\) in a binary tree and \(Z\) is a fixed constant, the algorithm practically achieves  <i>logarithmic bandwidth</i> \(O(\log N)\), even in the worst-case, for any practical block size. This surpassed the previous standard of \(O(\frac{\log^2 N}{\log(\log N)})\) for small-client-storage oblivious RAMs, a cuckoo hashing scheme set forth by Kushilevitz et al. in  <a href="#citeproc_bib_item_9">[9]</a>. We only mention asymptotic results for now, as experimental results were only found for a construction in section  <a href="#org6d35cc6">Recursive Path ORAM</a>. 
</p>

 <p>
On a related note, the failure probability of this algorithm is solely determined by potential path collisions when updating \(\pi\). If too many blocks are mapped to the same path, then a fixed-size stash \(S\) can potentially overflow. Accordingly, Stefanov et al.  <a href="#citeproc_bib_item_8">[8]</a> constructed a novel proof structure characterizing the likelihood of stash overflow during the Path ORAM algorithm – this proof structure later became popular among other tree-based ORAM constructions. They were able to show that even at large security parameters (\(\lambda = 128\)) a stash of 147 blocks was sufficient for bucket size \(Z=4\). Notably, stash size is purely a function of \(Z, \lambda\), with no dependence on memory size \(N\). This made Path ORAM uniquely scalable to large memory sizes even with fixed client-side storage constraints. In practice, the paper experimentally found \(Z=4\) to be minimum required bucket parameter in order to prevent excessive stash accumulation, while stating that  <i>theoretically</i> \(Z=5\) might be a better candidate.
</p>

 <p>
There is also the benefit of  <i>usability</i> – Path ORAM, despite being a hyper-efficient algorithm, is significantly easy to describe and construct when compared to older oblivious sorting and cuckoo hashing schemes (like the one in section  <a href="#org9532a5c">Square Root ORAM</a>). This made ORAM a much more  <i>feasible</i> tool for not just cryptographers, but engineers and system architects to design secure systems.
</p>
</div>
</div>
</div>
 <div id="outline-container-org6d35cc6" class="outline-3">
 <h3 id="org6d35cc6">Recursive Path ORAM</h3>
 <div class="outline-text-3" id="text-org6d35cc6">
 <p>
Despite the stash being invariant of memory size, the mapping table \(\pi\) inherently must be. For large \(N\), it becomes infeasible to store this on a client with limited storage.
</p>

 <p>
Path ORAM borrows from ideas in  <a href="#citeproc_bib_item_3">[3]</a> and  <a href="#citeproc_bib_item_7">[7]</a>, specifically with regards to moving mapping tables off-client via recursive ORAMs. To put it informally:
</p>

 <blockquote>
 <p>
The 'big' ORAM's mapping table becomes an ORAM. Then that ORAM's mapping table becomes an ORAM. And then  <i>that</i> ORAM's mapping table…
</p>
</blockquote>

 <p>
This procedure keeps going until a client requires suitably low amounts of storage to initiate an ORAM access. While this seems like it would pollute bandwidth, as indicated by a higher worst-case bandwidth \(O(\log^2 N)\), in practice Recursive Path ORAM still achieves R/W bandwidth of \(O(\log N)\). For a given access sequence \(\vec{y}\), this was experimentally determined by  <a href="#citeproc_bib_item_4">[4]</a> to be approximately 160x the bandwidth of simply using an insecure memory. While this isn't directly comparable to the results on older schemes collected by  <a href="#citeproc_bib_item_3">[3]</a>, which were done using 64KB block sizes, there is a massive order of magnitude difference compared to older ORAM schemes. While this is less performant than the SSS construction in the same work, the relative simplicity and better client memory efficiency of Path ORAM made it significantly easier to iterate upon and implement in hardware-centric applications.
</p>
</div>
</div>
 <div id="outline-container-org0f7cb8b" class="outline-3">
 <h3 id="org0f7cb8b">Ring ORAM</h3>
 <div class="outline-text-3" id="text-org0f7cb8b">
 <p>
Ring ORAM, like Recursive Path ORAM, isn't so much a new algorithm altogether so much as it is a set of improvements to Path ORAM proposed by Ren et al.  <a href="#citeproc_bib_item_4">[4]</a>.
</p>

 <p>
Ring ORAM's first goal was to reduce the amount of  <i>online bandwidth</i>, which we can understand to be the required (non-amortized) bandwidth per access. To do so, they store additional encrypted metadata with each bucket containing a  <i>permutation</i> – e.g which block is stored in each position of the bucket. This means that instead of loading every bucket along the path \(P(leaf_i)\), two accesses are made per bucket - one to query the metadata, and one to load the appropriate block afterwards.
</p>

 <p>
A corollary of this approach is that we can no longer have an in-place eviction, as we'd be evicting many blocks out of each bucket prematurely (counter to the goal of reducing per-access bandwidth). Like the construction in section  <a href="#org9532a5c">Square Root ORAM</a>, Ring ORAM opts to periodically evict blocks in its lookaside storage (in this case, the stash instead of the shelter) back into the primary storage (buckets) and reshuffle when doing so.
</p>

 <p>
Between the initial access and the reshuffle, the block is  <i>invalidated</i> in the memory. To ensure we still read some block on every path traversion, a dummy block is read out in any case that the requested block address is not found in a bucket. To facilitate the dummy-pulling while employing a periodic eviction strategy, Ring ORAM makes a couple of modifications to Path ORAM's data model.
</p>

 <ol class="org-ol"> <li>Ring ORAM will store \(Z\) real blocks and \(S\) dummy blocks per bucket after each reshuffle. We retain the standard definition of a dummy block.</li>
 <li>An eviction is carried out  <i>either</i> every \(A\) accesses (a tunable parameter) or when a single bucket is accessed \(S\) times without eviction (this is tracked via a counter). This latter condition is in place so that an argument regarding a block  <i>having</i> to be real cannot be made per the pigeonhole principle prior to a bucket eviction.</li>
</ol> <p>
With the above changes in place, we can make an additional optimization via an  <i>XOR</i>. Since the requested block can only be found  <i>once</i> on the traversion path, we can guarantee that we will either pull out \(L+1\) dummy blocks or \(L\) dummies and the requested block. Since dummies are assumed to always have a plaintext known to the client, we can assume that provided the nonces associated with each dummy, the client can generate the expected ciphertext for each dummy.
</p>

 <p>
By returning an  <code>XOR</code> of all the pulled data blocks from the memory, and then performing an  <code>XOR</code> on the client side between this block and each of the locally generated dummy ciphertexts, we can still pull out the expected ciphertext of the  <i>real</i> block. In doing so, the memory will only have to send back  <i>one</i> effective data block on each path traversal, regardless of tree depth. In short, the per-access bandwidth becomes  <i>negligibly conditional on bucket size and tree depth</i>. This is perhaps the biggest contribution of Ring ORAM – an online access bandwidth that is  <i>approximately identical</i> to the expected bandwidth with insecure memory access.
</p>

 <p>
As a whole, the improvements proposed by Ring ORAM reduce the bandwidth required for a single access  <i>without</i> eviction from the tree to constant overhead. However since evictions are still required, the asymptotic amortized bandwidth (e.g average bandwidth over many accesses) is still \(O(\log (\frac{N}{A}))\), a logarithmic bound.
</p>

 <p>
That said, this still translates to a practical performance improvement. Testing results in  <a href="#citeproc_bib_item_4">[4]</a> showed that under the same 4KB testing scheme referred to in section  <a href="#org9b86754">Performance and Usability Implications</a>, Ring ORAM with the XOR trick were able to bring overall bandwidth to only  <i>60x</i> the bandwidth required for insecure memory accesses. This is a little over \(\frac{1}{3}\) the bandwidth required for a recursive Path ORAM in the same scenario.
</p>
</div>
</div>
</div>
 <div id="outline-container-orgcdbe140" class="outline-2">
 <h2 id="orgcdbe140">Related Research</h2>
 <div class="outline-text-2" id="text-orgcdbe140">
 <p>
In this section we, at a high level, discuss a variety of adjacent research endeavors that can leverage ORAM mechanisms as a facet of their system architecture or attempt to iterate on the construction of modern ORAM schemes.
</p>
</div>
 <div id="outline-container-org5f3c112" class="outline-3">
 <h3 id="org5f3c112">Applications</h3>
 <div class="outline-text-3" id="text-org5f3c112">
 <p>
While most of this primer discusses ORAM in a single-party context, the reality is that remote computation is rarely a single-client system anymore. In scenarios where a group of trusted users require concurrent and secure access to some remote memory, many single-client ORAM schemes can encounter  <i>coherency issues</i>. As a minimum motivating example, we can look at the basic Path ORAM construction - there are possible scenarios where if you have 2 clients with stashes that contain real blocks, client A may unable to access a block that is temporarily held in client B's stash.
</p>

 <p>
Anticipating this coherency problem, Goodrich et al. proposed a stateless ORAM mechanism for secure group data access  <a href="#citeproc_bib_item_10">[10]</a>. Importantly, this mechanism requires that any client state be downloaded and re-uploaded to the remote when a client performs a memory access. As a result, any state maintained by the ORAM protocol should be  <i>lean</i> to prevent bandwidth pollution. While constructions like SSS  <a href="#citeproc_bib_item_3">[3]</a> may not be well-suited to this landscape due to high client storage requirements, a construction like Recursive Path ORAM  <a href="#citeproc_bib_item_8">[8]</a> is better suited to this task. Its position map and stash can be maintained with only \(O(\log N)\) client storage and \(O(\log^2 N)\) bandwidth.
</p>

 <p>
Another field with direct applications for ORAM is in secure processors. While there is a plethora of ongoing research regarding encrypted compute in enclaves and secure processors, we focus more on the utility of ORAM to mask program locality and memory access behavior. These architectures, approach memory interactions with a similar lens to the original Goldreich-Ostrovsky ORAM constructions  <a href="#citeproc_bib_item_2">[2]</a>, trying to mask information about the underlying program instructions that a secure processor is fetching out of RAM.
</p>

 <p>
Simpler, more bandwidth-efficient constructions like Path ORAM that eschew in-memory sorts and shuffles lend themselves to better hardware representations. Hardware design generally benefits from eschewing large fully-associative structures and maintaining small fixed-bound iterations for anything repetitive. Both of these criteria are fulfilled by the Path ORAM model  <a href="#citeproc_bib_item_8">[8]</a>. This is best embodied by some of the work done by Fletcher et al. in  <a href="#citeproc_bib_item_11">[11]</a>,  <a href="#citeproc_bib_item_12">[12]</a>. 
</p>
</div>
</div>
 <div id="outline-container-orga969eb4" class="outline-3">
 <h3 id="orga969eb4">Theoretic</h3>
 <div class="outline-text-3" id="text-orga969eb4">
 <p>
We mentioned in section  <a href="#org920cb7a">Drawbacks</a> that  <a href="#citeproc_bib_item_2">[2]</a> proved a lower bound on bandwidth of the form \(\Omega(t \log t)\) for \(t\) accesses. Notably, this bound is invariant of client memory size or block size.
</p>

 <p>
However, this bound assumed an ORAM scheme  <i>based on oblivious sorting</i>. Later papers like  <a href="#citeproc_bib_item_13">[13]</a> showed that if a construction eschews this mechanism (like Path/Ring ORAM), a system can bring this bandwidth bound down to \(\Omega(t\log(\frac{tr}{c}))\) for \(t\) accesses, where \(r\) is the size of each block in an N-block memory and \(c\) is the size of client memory. This conclusion is much more in line with the optimizations that schemes like SSS in  <a href="#citeproc_bib_item_3">[3]</a> make for large client memory sizes, and presents a harder target for novel ORAM schemes to pursue.
</p>

 <p>
At the same time, many newer ORAM constructions were published after the Path and Ring ORAM papers that continued to optimize bandwidth under low client-side storage constraints, or relaxed other assumptions underpinning amortization of these schemes. A good example is OptORAMa  <a href="#citeproc_bib_item_14">[14]</a>, which achieved \(O(\log N)\) access bandwidth overhead for small memory word sizes and limited client-side memory. This improves upon the \(O(\log^2 N)\) overhead of prior constructions like Path ORAM in similar regimes. This improvement stemmed from Intersperse-based shuffling and more efficient hashing schemes (BigHT, SmallHT) used during reshuffling.
</p>

 <p>
While modern constructions discussed in this primer are efficient enough to be used in real-world systems, future research will need to approach ORAM not just as a generic algorithmic tool but as something specialized to specific use cases. For example, future ORAM schemes with minimized eviction and shuffling and small stash sizes are likely to be more suitable for hardware implementation – a frontier that sees more relevance than ever in today's compute landscape.
</p>
</div>
</div>
</div>
 <div id="outline-container-org8e896d0" class="outline-2">
 <h2 id="org8e896d0">References</h2>
 <div class="outline-text-2" id="text-org8e896d0">
 <style>.csl-left-margin{float: left; padding-right: 0em;}
 .csl-right-inline{margin: 0 0 0 2em;}</style> <div class="csl-bib-body">
   <div class="csl-entry"> <a id="citeproc_bib_item_1"></a>
     <div class="csl-left-margin">[1]</div> <div class="csl-right-inline">N. Pippenger and M. J. Fischer, “Relations among complexity measures,”  <i>J. acm</i>, vol. 26, no. 2, pp. 361–381, Apr. 1979, doi:  <a href="https://doi.org/10.1145/322123.322138">10.1145/322123.322138</a>.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_2"></a>
     <div class="csl-left-margin">[2]</div> <div class="csl-right-inline">O. Goldreich and R. Ostrovsky, “Software protection and simulation on oblivious rams,”  <i>J. acm</i>, vol. 43, no. 3, pp. 431–473, May 1996, doi:  <a href="https://doi.org/10.1145/233551.233553">10.1145/233551.233553</a>.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_3"></a>
     <div class="csl-left-margin">[3]</div> <div class="csl-right-inline">E. Stefanov, E. Shi, and D. Song, “Towards practical oblivious ram,” Feb. 2012. Available:  <a href="https://www.ndss-symposium.org/ndss2012/towards-practical-oblivious-ram/">https://www.ndss-symposium.org/ndss2012/towards-practical-oblivious-ram/</a></div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_4"></a>
     <div class="csl-left-margin">[4]</div> <div class="csl-right-inline">L. Ren  <i>et al.</i>, “Constants count: Practical improvements to oblivious ram,” in  <i>Proceedings of the 24th usenix conference on security symposium</i>, in Sec’15. Washington, D.C.: USENIX Association, 2015, pp. 415–430.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_5"></a>
     <div class="csl-left-margin">[5]</div> <div class="csl-right-inline">B. Pinkas and T. Reinman, “Oblivious ram revisited,” in  <i>Proceedings of the 30th annual conference on advances in cryptology</i>, in Crypto’10. Santa Barbara, CA, USA: Springer-Verlag, 2010, pp. 502–519.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_6"></a>
     <div class="csl-left-margin">[6]</div> <div class="csl-right-inline">M. T. Goodrich and M. Mitzenmacher, “Mapreduce parallel cuckoo hashing and oblivious RAM simulations,”  <i>Corr</i>, vol. abs/1007.1259, 2010, Available:  <a href="http://arxiv.org/abs/1007.1259">http://arxiv.org/abs/1007.1259</a></div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_7"></a>
     <div class="csl-left-margin">[7]</div> <div class="csl-right-inline">E. Shi, T. .-H. H. Chan, E. Stefanov, and M. Li, “Oblivious ram with o((logn)3) worst-case cost,” in  <i>Advances in cryptology – asiacrypt 2011</i>, D. H. Lee and X. Wang, Eds., Berlin, Heidelberg: Springer Berlin Heidelberg, 2011, pp. 197–214.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_8"></a>
     <div class="csl-left-margin">[8]</div> <div class="csl-right-inline">E. Stefanov  <i>et al.</i>, “Path oram: An extremely simple oblivious ram protocol,”  <i>J. acm</i>, vol. 65, no. 4, Apr. 2018, doi:  <a href="https://doi.org/10.1145/3177872">10.1145/3177872</a>.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_9"></a>
     <div class="csl-left-margin">[9]</div> <div class="csl-right-inline">E. Kushilevitz, S. Lu, and R. Ostrovsky, “On the (in)security of hash-based oblivious ram and a new balancing scheme,” in  <i>Proceedings of the twenty-third annual acm-siam symposium on discrete algorithms</i>, in Soda ’12. Kyoto, Japan: Society for Industrial and Applied Mathematics, 2012, pp. 143–156.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_10"></a>
     <div class="csl-left-margin">[10]</div> <div class="csl-right-inline">M. T. Goodrich, M. Mitzenmacher, O. Ohrimenko, and R. Tamassia, “Privacy-preserving group data access via stateless oblivious ram simulation,” in  <i>Proceedings of the twenty-third annual acm-siam symposium on discrete algorithms</i>, in Soda ’12. Kyoto, Japan: Society for Industrial and Applied Mathematics, 2012, pp. 157–167.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_11"></a>
     <div class="csl-left-margin">[11]</div> <div class="csl-right-inline">C. W. Fletcher, M. v. Dijk, and S. Devadas, “A secure processor architecture for encrypted computation on untrusted programs,” in  <i>Proceedings of the seventh acm workshop on scalable trusted computing</i>, in Stc ’12. Raleigh, North Carolina, USA: Association for Computing Machinery, 2012, pp. 3–8. doi:  <a href="https://doi.org/10.1145/2382536.2382540">10.1145/2382536.2382540</a>.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_12"></a>
     <div class="csl-left-margin">[12]</div> <div class="csl-right-inline">L. Ren, X. Yu, C. W. Fletcher, M. van Dijk, and S. Devadas, “Design space exploration and optimization of path oblivious ram in secure processors,”  <i>Sigarch comput. archit. news</i>, vol. 41, no. 3, pp. 571–582, Jun. 2013, doi:  <a href="https://doi.org/10.1145/2508148.2485971">10.1145/2508148.2485971</a>.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_13"></a>
     <div class="csl-left-margin">[13]</div> <div class="csl-right-inline">K. G. Larsen and J. B. Nielsen, “Yes, there is an oblivious ram lower bound!,” in  <i>Advances in cryptology – crypto 2018: 38th annual international cryptology conference, santa barbara, ca, usa, august 19–23, 2018, proceedings, part ii</i>, Santa Barbara, CA, USA: Springer-Verlag, 2018, pp. 523–542. doi:  <a href="https://doi.org/10.1007/978-3-319-96881-0_18">10.1007/978-3-319-96881-0_18</a>.</div>
  </div>
   <div class="csl-entry"> <a id="citeproc_bib_item_14"></a>
     <div class="csl-left-margin">[14]</div> <div class="csl-right-inline">G. Asharov, I. Komargodski, W.-K. Lin, K. Nayak, E. Peserico, and E. Shi, “Optorama: Optimal oblivious ram,”  <i>J. acm</i>, vol. 70, no. 1, Dec. 2022, doi:  <a href="https://doi.org/10.1145/3566049">10.1145/3566049</a>.</div>
  </div>
</div>
</div>
</div>
 <div id="footnotes">
 <h2 class="footnotes">Footnotes: </h2>
 <div id="text-footnotes">

 <div class="footdef"> <sup> <a id="fn.ackshually" class="footnum" href="#fnr.ackshually" role="doc-backlink">1</a></sup> <div class="footpara" role="doc-footnote"> <p class="footpara">
This assumes that the input program is large enough that \(\log t > 1\).
</p></div></div>


</div>
</div></div>]]></description>
  <link>https://pradyun.net/blog/intro_to_oram.html</link>
  <guid isPermaLink="false">https://pradyun.net/blog/intro_to_oram.html</guid>
  <pubDate>Sat, 10 May 2025 00:00:00 +0000</pubDate>
</item>
<item>
  <title>Semi-Objectively Judging (Indian) Movies</title>
  <description><![CDATA[<div id="content" class="content">

 <p>
I tend to be rather outspoken about my opinion on any movie, but particularly Indian movies. Indian movies (primarily Telugu) have been my primary source of media consumpion for most of my life. As a result, it's one of the few things that I feel comfortable forming relatively well-based opinions about. It's 2AM and I'm waiting for my bedsheets to dry, so here's a dump on how I like to judge movies.
</p>

 <p>
Indian movies tend to suffer from either hyperinflated or unfairly-deflated ratings when viewed from the lens of someone more accustomed to western cinema (as their primary source of entertainment). I'm sure we've all encountered some occurence of 'well what moron would believe that you could hit a guy and he would fly back 6 feet'. Aside from the over-the-top action scenes, there tends to be criticism on the primarily "hero-centric" plot of many "masala" movies, and the rather formulaic storylines of "love at first sight" and an almost  <a href="https://en.wikipedia.org/wiki/Tsundere">tsundere</a> romance plot from the perspective of the heroine. Maybe I'm just desensitized to it, but I do think that those are facets inherent to the art form, and I tend to not factor it too much into how I judge a movie (unless the handling of either the action or romance in a movie is especially egregious).
</p>

 <p>
However as it is with most debates on "whether something is good", there tends to be a decent amount of personal opinion involved, so  <i>some</i> methodology of analysis generally helps to cast discussion over the quality of a movie into a more objective (or healthy) discussion. Of course there's always the option to simply  <i>not</i> discuss what you think of the quality of the movie, but what's the fun in that?
</p>
 <div id="outline-container-orgd9e418e" class="outline-2">
 <h2 id="orgd9e418e">Detailing a  Somewhat Sane Methodology</h2>
 <div class="outline-text-2" id="text-orgd9e418e">
 <p>
First things first, I like to separate the  <i>vision</i> of a movie from its  <i>execution</i>. Poor execution might be the result of low budget, uncooperative cast/crew, and all other manner of factors. But a poor  <i>vision</i> can make for a bad movie from start to finish, regardless of how well you dress it up. This is the difference between a movie like  <i>RRR</i> and  <i>1 Nenokkadine</i>. RRR tends to be publicly touted as "one of India's cinematic masterpieces", but I'd argue it's a perfect example of perfect  <i>execution</i> of a poor  <i>vision</i>.
</p>

 <p>
The best way to see that might be to imagine what RRR would be if we  <i>remove</i> some of the "hallmark" scenes. NTR's clash with Ramcharan with all the animals and water hoses, the final fight with the two against the King, and some of the other more "extravagant" fight scenes. I feel like if you do this (which may be unfair to the film genre, by some people's opinion) then the movie immediately begins to fall flat - the plot is rather lackluster, and the movie overall attempts to justify itself by doubling down on nationalist sentiment. Overall, the vision of RRR is likely a "patriotic  <i>spectacle</i>", and it achieves that well, but I don't know if it makes for a  <i>good movie</i>, at least by my preference for either good plot or abundant comedy. Definitely makes for good press and eye-candy though.
</p>

 <p>
By contrast, 1 Nenokkadine is a movie where some of the execution is a  <i>disservice</i> to the vision. The plot overall is an amazing concept - a schizophrenic musician kills off his parents' murderers, only to realize he hallucinated the murders. The rest of the movie unfolds as he struggles to differentiate what parts of an elaborate conspiracy are fact or fiction. Mahesh Babu puts forward one of his career best performances in the movie, playing the tormented musician. However, the impact of the movie overall is dulled by two big things - poor pacing and plot direction in the latter half of the movie, as well as an overall lack of refinement in the way villain encounters are handled.
</p>

 <p>
Some exceptions are there. The pre-interval scene shows our protagonist kill a primary antagonist thinking it to be a hallucination, only to find out he has killed the culprit in real life.  <i>Peak</i> cinema when shown in context. But by contrast, the film is chock-full of dead space like half-baked yacht fight and bridge fight scenes. They act as a filler that don't satisfy the critic, nor the masala-loving filmgoer. The combination of the two factors above are, at least in my opinion, why the movie didn't fare well in the box office. If I were to remake a movie, 1 Nenokkadine would definitely be a top choice on my list of candidates.
</p>

 <p>
As a side note, something that's been irking me about newer movies is a lack of continuous character development or screenplay. One of the bigger examples of this for me recently has been Kamal Haasan's portrayal of Vikram. He has a great performance and the fight scenes are beautiful - the main reasons why the movie fared so well - but his motivation as a character effectively unravels towards the end. We witness the narrative from the perspective of a man attempting to avenge the unjust murder of his son, only to encounter a climax monologue about a drug-free society and no mention of the dead son. The disconnect, especially since it was towards the end, left a sour taste in my mouth, and gave me a negative final opinion on the movie.
</p>

 <p>
Vidaamuyarchi suffered from an opposite problem - it spends about 40 minutes on exposition of why a husband and wife met, married, grew apart, then divorced. This time is all spent incessantly cutting between different flashbacks, in a seemingly random order. I feel that giving these flashbacks in a more orderly manner - for example, recounting the circumstances of their relationship to a divorce lawyer as a conversation - would have made for a much more linear narrative, cutting this exposition to only 25 or 30 minutes. The remaining time could have been used to pace the final stages of the movie better, which felt extremely rushed in the chosen screenplay.
</p>

 <p>
This leads me to believe that the  <i>writing</i> and  <i>screenplay</i> for movies has been weakening, while the idea and plots still have some originality or quality to them. Though that's just my opinion as a tired amateur - I'm sure experts have done much more in-depth analysis of these things.
</p>
</div>
</div>
</div>]]></description>
  <link>https://pradyun.net/blog/film_evaluations.html</link>
  <guid isPermaLink="false">https://pradyun.net/blog/film_evaluations.html</guid>
  <pubDate>Thu, 06 Feb 2025 00:00:00 +0000</pubDate>
</item>
<item>
  <title>[In]sane Default Formatting</title>
  <description><![CDATA[<div id="content" class="content">

 <p>
For some inane reason, nearly every editor and formatter out there defaults to 4-space indents. Which is fine if you're a psychopath, I suppose? I mean look at the difference between the same snippet with 4 spaces and 2 spaces.
</p>

 <div class="org-src-container">
 <pre class="src src-rust">for stream in listener.incoming() {
    let mut newconn: MyConn;

    match stream {
        Ok(s) => {
            newconn = MyConn{
                stream: s,
                name: String::new()
            };
        },

        Err(ref e)
            if e.kind() == ErrorKind::WouldBlock
                => {break;},

        Err(e) => panic!("conn err: {e}"),
    }
</pre>
</div>

 <div class="org-src-container">
 <pre class="src src-rust">for stream in listener.incoming() {
  let mut newconn: MyConn;

  match stream {
    Ok(s) => {
      newconn = MyConn{
        stream: s,
        name: String::new()
      };
    },

    Err(ref e)
      if e.kind() == ErrorKind::WouldBlock
        => {break;},

    Err(e) => panic!("conn err: {e}"),
  }
</pre>
</div>

 <p>
Night and day.
</p>

 <p>
I guess it's fine if the tab width is  <i>defaulted</i> to four spaces and I could override it to the  <i>obviously better option</i> of two spaces without putting my leg behind my neck, but of course not. Case in point, this is the snippet of my Emacs configuration required just to make  <i>Verilog</i> indentation default to 2 spaces.
</p>

 <div class="org-src-container">
 <pre class="src src-emacs-lisp">( <span style="font-weight: bold;">use-package</span> verilog-ts-mode
   <span style="font-weight: bold;">:defer</span> t
   <span style="font-weight: bold;">:custom</span>
  (verilog-case-indent 2)
  (verilog-cexp-indent 2)
  (verilog-indent-level 2)
  (verilog-indent-level-module 2)
  (verilog-indent-level-directive 2)
  (verilog-indent-level-declaration 2)
  (verilog-ts-indent-level 2))
</pre>
</div>

 <p>
Sorry about the lack of syntax highlighting in the above snippets. Website is still getting sorted out/setup, I'll try to get something sorted.
</p>

 <p>
Regardless. Either editors and tooling need to default to the superior standard of 2 space indents or they need to make my life easier when changing it to 2 spaces. However, I doubt either of those will happen before the heat death of the universe.
</p>

 <blockquote>
 <p>
 <b>Note:</b> After discussion with a friend (and some casual scripting last weekend) it comes to mind that Python would look utterly horrid with a 2-space indent structure. I do still think that the general viewpoint of this post is accurate, albeit with the caveat that it only holds for  <i>sane</i> languages that utilize some sort of delimiter to indicate discrete semantic sections. Python using whitespace to describe code structure has always left an odd impression on me, and it's hard to generalize the formatting recommendations made here to it as such.
</p>
</blockquote>
</div>]]></description>
  <link>https://pradyun.net/notes/fourspaces.html</link>
  <guid isPermaLink="false">https://pradyun.net/notes/fourspaces.html</guid>
  <pubDate>Thu, 30 Jan 2025 00:00:00 +0000</pubDate>
</item>
<item>
  <title>Standardized Typography</title>
  <description><![CDATA[<div id="content" class="content">

 <p>
Over the years one of my pet peeves has been seeing different fonts across my system. Feels jarring almost, especially considering the lengths I went to to make my system colors  <a href="https://github.com/pradyungn/Mountain">coherent</a>. This has led me to attempt to standardize typography across the different applications and documents that I maintain.
</p>

 <p>
Usually I like to use Outfit for headlines and title text, Crimson Text for any serif/prose, and a custom build of Iosevka (Myosevka) for any code. Myosevka is of my own selection, and is simply a low-weight condensed version of Iosevka.
</p>

 <p>
Unfortunately that standardization can break, this website (for now) being an example of this. At the same time I have been attempting to get PragmataPro to play nice with Emacs for a long time to no avail - something to do with Iosevka having sufficient font hinting, but PragmataPro not? Regardless, despite the font looking  <i>great</i>, it looks like a blurry turd in Emacs. Thus PragmataPro lives in my terminal, while Myosevka lives on an interim basis.
</p>

 <p>
This is one of those things where I think it's a "once you start noticing it, it gets hard to ignore it" situation. I used to be rather permissive with fonts, but nowadays writing code in any non-condensed font gives me the heebie-jeebies. Same with writing documents in a sans-serif font (god forbid), with the exception of Google Docs (a necessary evil in life).
</p>

 <blockquote>
 <p>
 <b>Note</b>: As early as  May 1st, I got a license of PragmataPro VF to play cleanly with Emacs on both Mac OS and Linux! TL;DR, font sizes are  <i>not</i> readable at the same intervals on PragmataPro that they are on Iosevka. Linux also has some issues with its default hinting settings not playing well with PragPro. Regardless, my default monospace font is now PragmataPro, with no salient exceptions.
</p>
</blockquote>
</div>]]></description>
  <link>https://pradyun.net/notes/typography.html</link>
  <guid isPermaLink="false">https://pradyun.net/notes/typography.html</guid>
  <pubDate>Tue, 28 Jan 2025 00:00:00 +0000</pubDate>
</item>
</channel>
</rss>
